AI MissionsHR AutomationupgradedEnterprise Autonomy

An AI Mission for Performance Reviews

TS
Trevor Solis · Lead AI Engineer, Missions
March 25, 2026

The oldest defect in the performance review is not bias or a bad template. It is memory — a year written from the six weeks a manager can still recall — and the fix is not software that scores people.

In the last week of the cycle, a manager sits down with eight blank review forms and a year they cannot actually remember. The first one writes itself, because that person shipped something difficult in October and the whole team is still talking about it. The second is harder. There is a strong feeling that the year went well, an equally strong feeling that something went sideways in the spring, and nothing in between that can be described with any specificity. So the manager does what managers have always done: reaches for what is reachable. Recent work, vivid work, work that happened to be visible in a meeting the manager was in. By the fourth form the pattern has hardened into method, and by the eighth the manager is writing sentences that are true in tone and unverifiable in fact, because the concrete instances that would have supported or refuted them have quietly fallen out of reach.

What comes out the other end looks like a record. It has dates on it, it goes into a file, it gets referenced in conversations about the person's future, and everyone involved treats it as an account of what happened. But it is not a record. It is a memory artifact — a reconstruction assembled under time pressure from whatever survived in one person's head, weighted heavily toward the end of the period and toward whatever was loud rather than whatever mattered. The gap between those two things, between a record and a reconstruction, is the single largest quality problem in performance management, and it is almost never the problem that gets solved. We keep redesigning the form.

Recency is not a bias to correct, it is the shape of the document

It is tempting to file this under bias, alongside all the other cognitive tendencies that get a paragraph in manager training and then no further attention. That framing understates it. Recency is not a distortion applied to an otherwise sound process; it is the structural condition under which the document is produced. A person's working year consists of a few hundred meaningful episodes, most of which were small at the time — a piece of unglamorous cleanup that prevented three months of pain, a difficult conversation handled well with someone who never mentioned it, a stretch in February where the person was carrying two roles and nobody wrote it down. A manager with a full job of their own has no mechanism for retaining these. They are not being careless. They are being human at a scale that human memory was never built for, and the review form asks them to pretend otherwise.

The consequences run in both directions, which is what makes it more than an inconvenience. The person whose best work happened early is systematically under-described, and worse, they can feel it happening without being able to name it — the review is not wrong about anything in particular, it is just thin, and thinness reads as an assessment. The person whose recent quarter was strong gets a document that generalises a good six weeks into a good year. And because reviews accumulate, the distortion compounds: next year's document is written partly against last year's, so a thin account becomes the baseline against which the following account is drafted. What began as a lapse of memory ends up as a durable organisational fact about a person, and nobody involved ever made a decision to record something untrue.

The honest managers know all of this. Ask any of them what they would need to write a better review and the answer is almost never a better rubric or a five-point scale instead of a four-point one. It is some version of: I need to remember what actually happened. That is a retrieval problem, and retrieval is exactly the sort of thing software has been genuinely good at for decades. It is also the only part of the performance review that software has any business touching.

The moment it starts scoring, it stops being useful

There is enormous commercial pressure to walk past the retrieval problem and go straight to the judgment, because judgment is what looks impressive in a demo. A system that surfaces the February everyone forgot is quietly valuable and hard to put on a slide, while a system that outputs a number next to a person's name looks like the future. And so the market fills with products that promise to rate, rank, calibrate, flag underperformance, or predict who is likely to leave — and increasingly with the instrumentation that such promises require, because you cannot score people without measuring them, and once you have decided to measure people you end up measuring the wrong things.

That instrumentation is worth naming as a category and rejecting as a category: keystroke counts, screen capture, activity and idle-time tracking, sentiment analysis run across employees' messages, aggregate productivity scores built from any of the above. The objection is not squeamishness. It is that these things corrupt the exercise they claim to serve. They measure presence rather than contribution, and the two have never correlated well in any job worth doing. They are trivially gamed the moment people learn what is being counted, which converts a workplace into a performance staged for a sensor. They collect enormous quantities of material about a person rather than about their work. And they produce conclusions whose provenance is a model's internal weighting rather than an event anyone can point to, so when the person asks why — and they will — no answer is available except that the system said so.

That last point is where this stops being a philosophical objection and becomes an operational one. A performance document is not just an artifact of management; it is something the organisation may one day have to stand behind, in front of the person it describes and in front of whoever else eventually asks. A sentence that says "in the March migration, you took over the rollback plan when the original owner was out, and it worked" can be stood behind, because the underlying event exists independently of anything that wrote the sentence. A number that says "3.2, performance risk" cannot be stood behind by anyone, including the person who generated it. Organisations tend to discover this asymmetry late, after the system is embedded, which is one reason the current wave of autonomy projects has such a poor survival rate — Gartner has predicted that over forty percent of agentic AI projects will be cancelled by the end of 2027, citing unclear business value and inadequate risk controls among the causes. In performance management the risk control is not a feature you add later. It is the entire design constraint, and it points in one direction: build the thing that helps a manager remember, and refuse to build the thing that decides.

A mission that recalls, and a human who assesses

What a well-constructed AI Mission does here is narrow and unglamorous, and that is the point. It works from the evidence an organisation already generates in the ordinary course of doing work — the things a person shipped, the tickets they closed, the documents they wrote, the reviews they gave, the projects they were assigned to — because those are records of work rather than surveillance of a worker, and they existed before anyone thought to write a review. It organises that evidence by period rather than by recency, so that February is as retrievable as November. It groups related episodes into themes a manager might not have connected. It cites, every time, so that each item the manager sees comes attached to the thing it came from and can be inspected, disputed, or discarded. And it stops there, handing the manager a well-organised body of evidence and an empty page, which is exactly the right division of labour: the machine has done recall, and the human does the assessment.

Everything downstream of that page belongs to the manager, and this has to be a design commitment rather than a policy sentence. The mission does not rate. It does not rank people against each other, calibrate distributions, or produce a summary judgment that a manager is invited to accept. It does not identify who should be promoted, paid more, or let go, and it plays no role in those decisions beyond making sure the human making them can see the whole year rather than the end of it. It infers nothing about anyone's characteristics, circumstances, or state of mind, because those are not what it is looking at and not what a review is about. On a platform like StudioX, this is what Human-in-the-Loop means in its strictest form — not a human approving a machine's conclusion, but a machine that never forms one, and it is a useful test to apply to anything in this category that the broader body of work on the autonomous enterprise describes as an agentic system for people operations. Ask what it would say if the human stopped reading. If the honest answer is "a rating," it is the wrong tool for this job regardless of how good the rating is.

The reframing worth carrying out of this is that a performance review should be judged the way you would judge a piece of research rather than the way you judge a form: by how much of the year it can actually point at. A review full of confident adjectives and no instances is a weak document no matter how favourable it is, because it describes a manager's impression rather than a person's work. A review that can show its work — here is March, here is the thing in June nobody mentioned at the time, here is the pattern across four episodes spread over ten months — is a strong one even when it is critical, because the person can engage with it, correct it, and recognise themselves in it. Software can move a great many reviews from the first category to the second, and that is a larger contribution than any scoring engine will ever make. The organisations that get this right will not be the ones whose systems form the sharpest opinions about their people. They will be the ones whose managers, sitting down in that last week of the cycle, finally have the whole year in front of them and have to do nothing harder, or more human, than read it and say what they think.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.