Technology & The Future

Tracked by Software, Judged by Nothing: How AI Surveillance Is Quietly Breaking the Workplace

A new study found that AI workplace monitoring produces worse outcomes than human oversight — and the reason turns out to be less about privacy than about something older and stranger.

Julian CrossApril 28, 202610 min read
Tracked by Software, Judged by Nothing: How AI Surveillance Is Quietly Breaking the Workplace

Picture a call center agent who spends four minutes on a call that the system flags as too long. The agent was helping an elderly customer who had gotten confused and frightened mid-call, spoke slowly, and needed an extra pass through the instructions. The agent handled it with patience and care. The customer thanked her and stayed. The algorithm logged a deviation. Her score dropped.

This is not a hypothetical designed to make a point about robots being heartless. It is a description of what algorithmic performance monitoring does structurally, at scale, every day. The tool is built to measure what is measurable: time, output, deviation from baseline. It has no way to register context. It cannot distinguish between a slow call that was necessary and a slow call that was wasted. It cannot know what a worker was thinking, or what a customer actually needed, or what a four-minute deviation saved in downstream complaints. The algorithm only has the number.

A study published in Nature's Communications Psychology[1] by researchers at Cornell[2] examined what happens when workplaces shift toward AI-driven monitoring, and the findings are striking in their specificity. AI surveillance, compared to human oversight, produced significantly more employee complaints, measurably worse performance outcomes, and higher rates of voluntary turnover. Workers did not perform worse because they were less afraid of being caught. They performed worse because something about being watched and judged by a system that cannot read context changes what work feels like from the inside — and that feeling has real behavioral consequences.

The study does not land as a simple story about privacy invasion or big-brother overreach, though those concerns are real and adjacent. It lands as something more specific: a finding about what humans need in order to function well at work, why algorithmic judgment frustrates that need in ways human oversight often does not, and what it means that companies are deploying these systems faster than they understand what they cost.

The Gap Between What Gets Measured and What Actually Happened

Workplace monitoring is not new. Factories have used time-and-motion studies[4] since Frederick Winslow Taylor was designing them in the early twentieth century. Supervisors have always watched workers. The difference now is scale, granularity, and the removal of a human in the feedback loop. Modern AI monitoring systems can track keystrokes, mouse movement, time between actions, call duration, response latency, idle periods, eye gaze via webcam, and dozens of other behavioral signals simultaneously, continuously, and without a person reviewing each data point in real time. The system surfaces its conclusions as scores, flags, or automated prompts. A manager may review the outputs periodically. The surveillance itself is constant and automated.

What this creates is a new kind of managerial relationship — one in which the entity doing the most active moment-to-moment evaluation of your work is not a human being. And this distinction turns out to matter more than employers initially assumed, because human judgment is not simply slower or less precise automated judgment. It is a categorically different thing. Human managers bring context, history, relationship, and the ability to ask follow-up questions. They can recognize when a bad metric reflects a good decision. They carry an implicit social contract with the workers they evaluate: I see you as a whole person, not just an output. Algorithms cannot extend that contract because they have no access to the information it requires.

“The algorithm only has the number.”

Researchers in organizational psychology have a framework for this problem that predates AI monitoring but maps cleanly onto it. Procedural justice — the perceived fairness of the process by which one is evaluated — is one of the most reliable predictors of workplace engagement, trust, and willingness to put in discretionary effort. When people believe the process is unfair, they do not simply get annoyed. They recalibrate how much they invest. They become more conservative in their behavior, less willing to take risks or go off-script in ways that might help customers or colleagues but would expose them to algorithmic penalties. They start optimizing for the metric rather than the outcome the metric was meant to approximate. And eventually, enough of them leave.

Gaming the Score Instead of Doing the Job

One of the subtler findings embedded in the Cornell research, and in adjacent work on algorithmic management, is that AI monitoring does not just measure behavior. It shapes it. Workers who know they are being tracked by an automated system that cannot read context often adapt their behavior to fit what the system rewards, rather than what the work actually requires. This is not cynicism or shirking. It is a rational response to an evaluation structure that penalizes contextual judgment.

The pattern shows up across industries. Delivery drivers who know their routes are algorithmically scored for speed may skip bathroom breaks or cut corners on vehicle safety checks. Customer service agents who are evaluated on call duration may rush customers who need more time, or close tickets as resolved before the underlying problem is actually fixed. Software engineers whose output is tracked by commit frequency may push smaller, more frequent updates that look productive rather than spending time on harder architectural problems that take longer to surface as visible work. In each case, the worker is not failing at the job. The worker is succeeding at the measurement. Those are increasingly different things.

“Workers start optimizing for the metric rather than the outcome the metric was meant to approximate.”

Economists call this Goodhart's Law[3]: when a measure becomes a target, it ceases to be a good measure. The phenomenon is old. What is new is the scale at which automated systems can enforce it, and the speed at which behavioral drift accumulates when the feedback loop runs continuously with no human in a position to notice when the score and the reality have diverged.

What Makes Being Watched by a Machine Feel Different

There is a psychological mechanism worth naming here, because it is not the obvious one. Most conversations about AI surveillance center on privacy — the sense that constant monitoring is intrusive, that being watched all the time is dehumanizing, that data about your behavior can be misused. These concerns are legitimate. But the Cornell study's finding is pointing at something adjacent and somewhat more specific: the problem is not just that workers are being watched. It is how algorithmic evaluation handles the interpretation of what it sees.

When a human manager judges your performance, there is an implicit negotiation available. You can explain. You can push back. You can provide context that changes the interpretation. The manager might still disagree with you, might still give you a bad review, but there is a social space in which your perspective can enter the conversation. You are not simply the output of a measurement. You are a person making a case. Algorithmic monitoring forecloses that negotiation by design. The system has already decided. There is no one to explain yourself to. The score is the verdict, and the verdict has no appeals process because the process that generated it has no mechanism for hearing you.

Psychologists who study organizational fairness have found that workers are remarkably tolerant of negative outcomes — being passed over for promotion, receiving critical feedback, being held to high standards — when they believe the process that produced those outcomes was fair and gave them a voice. The outcome matters less than the process. Algorithmic monitoring tends to produce the opposite experience: the outcome feels arbitrary because the process is opaque, and the opacity feels like an insult on top of an injury. You did not just get a bad score. You got a bad score from something that did not even try to understand what you were doing.

The Quit Rate Is Telling You Something

Higher voluntary turnover under AI monitoring is, at first glance, exactly what you would expect if workers simply disliked surveillance. But the Cornell findings are more specific than that: the increase in quit rates was not uniform across all types of monitoring. It was more pronounced under algorithmic management than under human oversight of comparable intensity. Workers being closely watched by human supervisors did not quit at the same elevated rates as workers being closely tracked by automated systems. This suggests the problem is not the watching itself. It is who — or what — is doing it.

This has significant implications for how employers should think about the cost of these systems. The business case for AI monitoring typically involves efficiency: you can track more workers with less managerial overhead, surface performance problems faster, and reduce the need for supervisory labor. These are real advantages, up to a point. But turnover is expensive. Depending on the role, replacing a worker costs between half and twice their annual salary when you account for recruiting, training, lost productivity during the transition, and the institutional knowledge that walks out the door. If AI monitoring is meaningfully increasing quit rates, the efficiency gains from reduced supervisory overhead may be partially or fully offset by elevated churn. The system that seemed cheaper to operate turns out to be more expensive to sustain.

“The system that seemed cheaper to operate turns out to be more expensive to sustain.”

And that is before accounting for what turnover does to the workers who stay. High quit rates signal something to the people who remain: this is a place where people do not last. That signal suppresses long-term investment in the job. Workers become less likely to build deep expertise, less likely to take initiative, less likely to identify with the organization's goals. They hedge. They treat the job as provisional. Which makes their performance less good, which the algorithm flags, which makes the environment worse, which encourages more people to leave.

The Design Problem Nobody Is Fixing

It would be convenient if this were a story about bad actors deploying surveillance maliciously, but most algorithmic management systems are not designed with cruelty as the goal. They are designed to solve a legitimate operational problem — how do you maintain visibility into workforce performance across large, distributed, or remote teams? — using tools that are genuinely powerful for the things those tools can measure. The problem is that the things those tools can measure are not the same as the things that make work good, and the gap between the two is where the damage accumulates.

The structural issue is that companies adopting these systems are making a choice that is partly technical and partly about what they believe work is. If work is fundamentally an input-output problem — effort goes in, measurable product comes out, variation is deviation — then algorithmic monitoring is a reasonable fit. If work involves judgment, context, relationships, timing, the exercise of discretion in unpredictable situations, or any activity whose value is not immediately visible in a metric, then automated evaluation systems will systematically misread it. Not occasionally. Structurally. Because context blindness is not a bug in these systems that better engineering will eventually fix. It is a feature of what automated measurement is.

Some companies are beginning to recognize this. There is growing interest in what researchers call human-in-the-loop monitoring — systems where algorithmic tools surface patterns and anomalies, but human managers make the interpretive judgments and retain the relationship with workers. This is more expensive than fully automated evaluation, but it preserves the social negotiation that makes performance feedback feel legitimate. It also keeps a human accountable for the evaluation, which changes the incentives around how that evaluation gets done.

What the Algorithm Cannot See Is the Job

The deeper issue the Cornell study surfaces is about what we lose when we allow automated systems to define what counts as performance. Most work, even routine work, contains a layer of invisible judgment: the choice to spend a few extra minutes on a frustrated customer, the decision to flag a problem before it escalates, the instinct to slow down when something feels wrong. These choices rarely show up cleanly in metrics. They often look, from the outside, like inefficiency. They are frequently what separates a functioning organization from a mediocre one.

When workers learn that these choices will be penalized rather than rewarded, they stop making them. Not because they do not care, but because caring has been made structurally costly. The algorithm cannot see good judgment. It can only see the deviation that good judgment sometimes requires. And so workers learn, gradually, to make fewer deviations. To stay inside the range the system will reward. To become, in other words, slightly less human in their work — not because the technology demanded it explicitly, but because the incentive structure quietly made it the rational thing to do. That is the kind of change that arrives as a slow drift in workplace culture before it announces itself as a problem in the quarterly numbers.

References

  1. Algorithmic versus human surveillance leads to lower perceptions of autonomy and increased resistance (nature.com)
    Publishes the Cornell study finding that AI workplace monitoring produces more complaints, worse performance, and higher turnover than human oversight.
  2. More complaints, worse performance when AI monitors work (news.cornell.edu)
    Provides the core research showing AI monitoring causes greater loss of autonomy than human oversight and increases complaints, reduced productivity, and quit rates.
  3. Goodhart's law (en.wikipedia.org)
    Defines Goodhart's Law—when a measure becomes a target, it ceases to be a good measure—explaining why workers optimize for metrics rather than actual job outcomes.
  4. Time and motion study (en.wikipedia.org)
    Establishes historical context that Frederick Winslow Taylor designed time-and-motion studies in the early twentieth century, predating modern AI monitoring.

About Julian Cross

Julian Cross writes about AI, automation, surveillance, digital identity, labor, human relationships with each other and automation, complex systems and attention — less about what new tools, studies and observations can do in theory than what they're already doing to how we work, spend, relate, and get measured. His work follows leads to the point where it stops being a product and starts being a condition.

More like this

AI Is Watching You Work. It's Also Deciding What Work Means.

AI Is Watching You Work. It's Also Deciding What Work Means.

Julian Cross 9 min
The Boss Isn't Watching You. The Algorithm Is. And It's Worse.

The Boss Isn't Watching You. The Algorithm Is. And It's Worse.

Julian Cross 9 min
Your Productivity Score Is Real. The Productivity It Measures Isn't.

Your Productivity Score Is Real. The Productivity It Measures Isn't.

Julian Cross 9 min