When AI Agents Report on Employees: Review Rights
Workflow agents quietly produce reports about people. The subjects deserve performance review grade rights: see the record, attach context, and get a human decision.

Last week, in an experiment run by an AI research lab and reported by Time, an AI agent managing a small San Francisco retail store recommended firing a human employee for chronic tardiness. Then it did it. There was a human in the loop, technically. The lab's staff let the process play out to see where it would go. Where it went was a person losing a job on the strength of an AI's account of her behavior.
Most companies read that story and file it under "we would never." Fair enough. Almost nobody is putting an agent in charge of terminations. But the interesting part of the story is not the firing. It is the file. An agent observed a person's work over weeks, formed a running judgment about her reliability, and produced the record that the decision rested on. That part is not an experiment; it is already happening at your company.
You are already in the file
Think about what agents actually do all day in a modern team. One watches the support queue and summarizes throughput. One tracks the sales pipeline and flags deals that have gone quiet. One preps the Monday ops review by pulling together who did what, what slipped, and what is blocked. One reads the incident channel and drafts the timeline of who responded when.
Every one of those outputs is, in part, a report about people. The pipeline summary says your deals are stale. The queue digest says your resolution times run long. The incident timeline says you acknowledged the page eleven minutes after it fired. The Monday brief says the migration slipped and your name is next to the milestone.
None of this was designed as employee monitoring, and that is exactly why it is spreading without anyone deciding it should. Nobody procured surveillance software. Nobody ran it past legal. The team just asked an agent to keep track of the work, and people are part of the work. The observation arrived as a byproduct, and byproducts do not get review processes.
Then the byproduct lands in a channel your manager reads. The agent was not evaluating you. But when the third consecutive weekly summary says your queue is the slow one, your manager is not making a fine distinction between "workflow telemetry" and "performance data." An impression forms, and impressions turn into ratings. At most companies, the person the summary describes has no idea it exists.
The agent's account is confident, specific, and missing everything that matters
The instinct here is to say the agent is just reporting facts, and facts are fair. Your tickets did take longer. The deal did go quiet. What is there to contest?
Plenty, and anyone who has ever been managed knows it. Your tickets took longer because you were handed the migration cases nobody else could touch. The deal went quiet because the buyer's company froze spending, and you know that from a phone call the agent never saw. You were late acknowledging the page because you were already on a call about the same outage, dealing with it before the alert fired.
Human managers get this wrong too. But a human manager's impression is understood by everyone, including the manager, to be an impression. It gets tested in conversation before it hardens. An agent's summary arrives wearing the costume of objectivity: timestamped, quantified, neatly formatted. It reads like measurement, but it is an interpretation built on partial data, and the parts it is missing are precisely the ones a person would lead with in their own defense.
An agent's summary arrives wearing the costume of objectivity: timestamped, quantified, neatly formatted. It reads like measurement.
There is a second problem stacked on the first. Agents compress. Whatever nuance existed in the raw activity gets flattened into a sentence like "response times in the eastern queue continue to lag." The flattening is the value; nobody wants a summary as long as the week it summarizes. But every compression is a judgment about what counts, and the person being compressed had no say in the criteria and often no knowledge that the compression exists.
If this sounds theoretical, run one test. Ask the people on any agent-assisted team whether they know what the agents in their workflows have said about their work in the past month. Then watch their faces. Most have never thought to ask. A few have thought about it a great deal and concluded there was no way to find out. Neither group is fine, and the second group has usually already updated its behavior in quiet, trust-eroding ways: doing less in channels the agent reads, keeping context in DMs, working to the metric instead of the job.
Companies already know how to solve this. They just haven't noticed it applies.
Here is the strange part. Organizations have spent decades building norms for exactly this situation, in a different room. It is called the performance review.
The norms are so standard we forget they were ever won. An evaluation of your work must be shown to you. It happens on a known schedule, not in secret. You get to respond, and your response attaches to the record. Vague impressions are supposed to be backed by examples. There is an escalation path if you think the assessment is unfair. HR did not adopt these rules out of sentimentality; every one exists because its absence produced disasters, grievances, lawsuits, and the departure of good people who could not get a hearing.
Now a new class of evaluator has entered the workflow. It observes continuously rather than quarterly. It reports weekly rather than annually. Its observations reach your manager without passing through you. And it operates under none of the norms, because nobody categorized it as evaluation. The handbook governs the appraisal your manager writes once a year and says nothing about the fifty-two agent summaries that shaped what she believes before she starts typing.
The fix is not to make agents stop noticing people. That is neither possible nor desirable; a report on the work that omitted the humans doing it would be useless. The fix is to extend the review-rights framework to cover any agent output that characterizes a person's work. Three rights carry most of the weight.
The right to see it. If an agent's summary names you or describes work attributable to you, you should be able to find and read it. Not through a records request, not by asking a manager who has to go digging. As a matter of workspace design, the observation should live somewhere you can already see. This is where the architecture question stops being abstract: when agents work in the same shared channels as the team, their reports about the team are visible to the team by default. The pipeline digest that mentions your stale deals appears in the channel you are in. Nothing about you is being said behind a login you do not have. When agents instead report through side tools and private dashboards, visibility becomes a feature someone has to remember to build, and mostly nobody does.
The right to attach context. Seeing the record matters only if you can answer it, in the same place, attached to the same thread, before it congeals into your manager's settled view. "The queue ran slow because I had the migration cases" belongs directly under the summary that says the queue ran slow, where every future reader, human or agent, will encounter both together. This is also the cheapest correction mechanism ever built: no meeting, no HR process, just a reply. And it improves the agent, because next week's summary gets drafted with this week's correction in its context.
The right to a human decision. No consequence for a person, formal or informal, should rest on an agent's characterization that the person has not seen and answered. The store experiment failed exactly here. A human approving an agent's firing recommendation is not oversight if the human's entire understanding of the situation came from the agent's own account. Review means someone weighed the agent's version against the person's version. If the person never got to give one, what happened was ratification, not review.
The line managers have to hold
There is a temptation, reading this, to treat it as an HR policy problem, write the addendum, move on. The policy matters, but the daily failure mode is smaller and closer to the ground: a manager reads six weeks of agent summaries, forms a view, and by the time any formal process begins, the view is the verdict. The paperwork just catches up.
So the operating rule for managers is simple to state and takes actual discipline to follow: an agent's observation about a person is a prompt for a conversation, never a conclusion. The summary that flags your report's slow queue is the beginning of "hey, what's going on with the eastern queue?" and not the quiet middle of a case file. Managers who use it the first way get the real benefit, which is finding problems while they are still small and blameless. Managers who use it the second way are building performance cases out of unrebutted hearsay from a witness that cannot be cross-examined.
And leaders should notice which way their tooling tilts. A workspace where agent reports are visible to everyone they describe makes the first path natural: the employee saw the same digest, expects the conversation, and often arrives with the missing context already posted. An architecture of private dashboards tilts everyone toward the second path, not out of malice but because secrecy is the default and defaults win.
The lab experiment ended with a firing that made the news precisely because it was so legible: one agent, one decision, one person. What is accumulating inside ordinary companies is less legible and bigger. Thousands of small agent observations, each individually reasonable, quietly assembling into pictures of people who have never seen the picture. You do not need new values to deal with this. You need the values you already wrote into the performance review chapter, applied to the evaluators you did not notice hiring.
The subjects of those reports are not asking for protection from being seen. They are asking for what anyone asks of a fair evaluation: show me the file, let me answer it, and have a human decide. That was reasonable when the evaluator was a person. It does not become less reasonable when the evaluator never sleeps.
FAQ
Are AI agents already monitoring employees?
Not by design, but in effect. Agents that track pipelines, queues, and incidents report on the people doing that work as a byproduct, and the summaries reach managers without passing through the people they describe.
If the agent only reports facts, what is there to contest?
The missing context. Tickets ran long because they were the hard migration cases; the deal went quiet because the buyer froze spending. The agent's account is confident, specific, and built on partial data.
What rights should employees have over agent reports about them?
Three: the right to see any agent output that characterizes their work, the right to attach context to it in the same place, and the right to a human decision that has weighed their version against the agent's.
How should managers use agent observations about their team?
As the opening of a conversation, never as a case file. A manager who acts on six weeks of unanswered agent summaries is ratifying hearsay from a witness that cannot be cross-examined.
How does a shared workspace change this?
When agents report in the channels the team works in, the people described see the reports by default and can answer in the thread. Private dashboards make secrecy the default instead.
Small hops. Big leap.
Every drafted follow-up, every synced table, every brief that writes itself is one small hop. Together they change how the team moves. Early access is open.
Get startedAuthor
Alex Shershebnev
Alex Shershebnev is a seasoned AI engineer and technology leader with over a decade of experience in AI, DevOps and MLOps. He is currently Lead DevRel at Zencoder, an AI coding assistant, and one of the founding members of the company, where he has spent the last two years shaping both the product and its developer ecosystem. Alex has spoken at more than 50 international conferences, establishing himself as a recognized voice on AI for coding, secure and responsible use of AI in software development, and the future of developer workflows.