A growing workforce now signs off on AI output it cannot question, change or refuse
By 21 September 2026, a database kept by the legal researcher Damien Charlotin had logged 2,046 court decisions worldwide in which a judge found that a filing relied on material invented by an AI system. The material included case law that does not exist, quotations nobody said and citations that lead nowhere. In 815 of those cases, the person who put their name to the document was a lawyer. Courts in the United States, Canada, Australia and elsewhere have fined those lawyers, referred them to regulators and named them in published judgments. The software that produced the citations appears in the record as background. The consequences fall on the human being who read the output, or failed to, and signed it.
The same thing happens outside the courtroom. In October 2025, Deloitte agreed to refund part of the A$440,000 fee for a report it had prepared for Australia's Department of Employment and Workplace Relations. Chris Rudge, an academic at the University of Sydney, had found fabricated references in it, along with a quotation attributed to a Federal Court judgment that the judgment did not contain. The corrected version disclosed that a generative AI model had been used in drafting. The public argument that followed was about who should have caught the errors.
Legal filings and consulting reports are the visible edge of something much larger. In banks, hospitals, insurers, software teams, newsrooms, call centres and government departments, a new kind of work has taken hold: reading what a system produced and deciding whether it is right. The person doing it is held responsible when errors get through. They usually cannot change how the system behaves, cannot see why it produced a particular answer, and often have no authority to reject the output outright. By most measures it is among the fastest-growing kinds of work in the world, though no labour statistic counts it, since it still has no settled name. Some organisations call it human-in-the-loop review. Most call it nothing at all and simply add it to somebody's existing job.
Hassan Abu Sheikh, Co-Founder of CNTXT AI, said the position becomes impossible when the organisation above it has not done its own work first. "When it has been done properly, the person reviewing output is catching the rare failure. When it has been skipped, that person has been made the entire quality control function of the organisation without being told so," he said.
The job is spreading through every industry without a title or a job description
PwC's 2026 Global AI Jobs Barometer analysed more than a billion job advertisements across 27 countries and territories. It found that entry-level roles in AI-exposed fields were seven times more likely than other entry-level roles to ask for skills traditionally expected of senior staff, such as judgement and leadership. Put simply, junior employees are increasingly asked to evaluate work they did not produce and may not yet be qualified to assess. The review burden never appears as its own line in those advertisements. It sits inside roles that still carry their old titles: paralegal, claims handler, analyst, junior developer, content moderator, customer service agent.
A role without a name has no job description, no training standard, no agreed workload and no professional body to argue for any of them. Kurt Muehmel, Head of AI Strategy at Dataiku, said the errors that matter most are the ones reviewers are least equipped to see. "The hardest errors from AI to catch are the ones that do not look like errors. There are certain obvious errors, and there are others which are more insidious, in the sense that they are more difficult to detect," he said.
Muehmel said generative AI has also undone the old idea of what a software defect is, which makes the reviewer's task harder. "We have heard that these models are non-deterministic, meaning that the same input does not always give the same output. It used to be that software was always deterministic, meaning that when you did one thing, only one thing should happen. The notion of what used to be faulty software, of what used to be bugs in software, is very different with generative AI, which is non-deterministic by design," he said. A reviewer who has checked a hundred outputs has learnt nothing certain about the hundred and first.
Carrying responsibility without control is the oldest recipe for strain at work
Occupational psychologists have known for decades which jobs wear people down fastest. The demand-control model developed by the sociologist Robert Karasek found that the most damaging combination is heavy responsibility paired with little say over how the work is done. The AI reviewer sits squarely in that quadrant, with a further complication: they are expected to vouch for reasoning they cannot see.
A second body of research, going back to studies of radar operators during the Second World War, describes what psychologists call vigilance decrement. People asked to watch for rare errors in a stream of mostly correct material become steadily worse at spotting them. The more reliable the system, the less often the reviewer finds anything, and the less often they find anything, the further their attention drifts. The researcher Madeleine Clare Elish gave the resulting position a name, the "moral crumple zone". It describes the human inside an automated system who absorbs the blame for a failure the way a car's crumple zone absorbs an impact, protecting the machine and the institution behind it.
Laxmi Nageswari, Chief AI Officer at Cloud Box Technologies, said the reviewer's formal position rarely matches the blame they receive. "When an enterprise AI solution makes errors, accountability is ultimately tied to the organisation deploying the system and not to the human reviewers, although this is not always stated directly or made visible. Human reviewers operate as the supervising arm of a workflow. They may have little to no power to influence the decisions taken by the AI system," she said.
She said the work itself takes a toll that employers need to watch. "Human oversight is critical, as AI can carry bias in certain decisions and lacks empathy and critical thinking. It is also important to assess the wellbeing of reviewers and operators, as the job can become exhausting. This is where a responsible leadership team comes in, to make the workflow better both for the operators and for the end users," she said.
A high approval rate usually means the reviewer has stopped reviewing
Most organisations measure review by volume: how many items cleared, how quickly, with how few delays. Nageswari said her company reads the same numbers the other way. "By definition, AI-generated outputs cannot be 100% accurate. With that in mind, we keep logs on every task executed, from the initial input to the output. We treat approvals without changes as warning signs, as the human reviewers might be fatigued, facing automation bias or rubber-stamping," she said.
She said what counts is how carefully each decision was made. "Every input and output carries a different level of complexity. By analysing the pattern of inputs, dwell time, complexity and outputs, we can assess whether outputs are being approved precisely and accurately. The aim is to understand how effective humans are at validating outputs, and not how many outputs were approved to begin with," she said. Dwell time, the seconds a reviewer spends on an item before deciding, is one of the few signals that separates a considered approval from a reflex.
Automation bias is now written into law. Article 14 of the EU AI Act, which covers human oversight of high-risk systems, requires that the people assigned to oversee them can stay aware of the tendency to rely automatically or excessively on a system's output. It also requires that they can decide not to use the system, disregard, override or reverse its output, and interrupt it altogether. Those are precisely the powers many reviewers lack in practice.
Nageswari said reviewers using her company's systems can refuse an output entirely. "Our system allows human reviewers to reject an AI-generated output straight away. This helps organisations avoid factually incorrect outputs that can cause reputational and financial damage. We have also put mechanisms in place that allow reviewers to approve, reject or request an edit to inputs and outputs, as well as flag responses for further review," she said. She added that rejection protects the system as well as the individual decision. "Rejecting outputs ensures the system does not succumb to automation bias, where people accept AI recommendations because they assume they are usually correct. It also helps with system tuning and prompt optimisation, so corrective measures can be taken against recurring errors."
Reviewers are asked to judge reasoning they are never shown
Visibility is the second missing piece. Nageswari said a reviewer needs far more than the answer on the screen. "Human reviewers must not trust AI blindly and must have access to audit logs, the initial input, the retrieved context, factual sources and confidence scores, among other data points. This helps them establish whether the AI system actually interpreted the inputs as intended, or whether it produced dense and ungrounded blocks of text," she said.
She said the depth of scrutiny should rise with the stakes. "Not all applications and use cases are subject to strict due diligence, and lower-risk uses may treat contextual evidence as sufficient. High-risk tasks and decisions should be based on audit logs, stronger controls, validation and complete traceability. This ultimately helps humans be the judge of whether the output is right or not," she said. The same records, she said, are what allow an organisation to work out afterwards where a failure began. "Detailed audit logs contain the initial input, the context data that was retrieved and interpreted, any changes humans made, the final decision timestamp and everything in between. They help organisations establish the timeline and the cause, how the process moved from input to output, and where liability sits."
Muehmel said individual checking, however careful, cannot catch problems that only show up across thousands of outputs. "That is where it is critically important to have good baseline data on what a good response is supposed to look like from the AI system, so that you can go through responses programmatically, beyond looking at each one with a thumbs up and a thumbs down, and actually assess whether you are seeing drift or an increasing error rate. It is in doing it in aggregate that you are going to identify those core errors," he said. He added that the individual still matters, within limits. "On an individual basis it is of course important for any one person to keep a clear eye on what is going out, especially if they are the human in the loop, and to use their judgement. The real solution long term is to test the responses against what is expected and to confirm whether they are accurate."
Review only protects anyone when the reviewer could have done the work themselves
Abu Sheikh said the review layer has too often been handed to the wrong people. "The review layer only functions as a safeguard when the person doing the reviewing has the expertise that would allow them to catch the error. The whole difficulty of this new job is that it has been handed to people who were never given that expertise, and they are then held to account as though they had it," he said.
He said an expert and a novice experience the same work in opposite ways. "A subject matter expert who used to do the task before AI can now come in with AI and multiply their productivity 20 times over, because they know what the work is supposed to look like and they can tell immediately when what has come back is wrong. Someone who opens a prompt window and produces a product requirements document has not become a product manager. Vibe coding does not make you an engineer. Just because I can make an AI generate blueprints for a rocket does not make me a rocket scientist," he said.
For the expert, he said, checking is part of the craft. "The review work is the cost of the productivity, and for the expert that trade is worth making, as they are still doing the task they were always doing, faster and at higher quality. The person for whom this becomes an impossible job is the person who was never doing the task in the first place and has been handed the output of a system they have no way of evaluating," he said.
He said blaming the tool lets the organisation escape its own decisions. "Anybody can come and say they did this while it is 100% wrong, and then they blame the tool, they blame the AI. If that were the case, woodworkers would blame hammers. Once an organisation accepts blaming the system as an explanation, it stops examining the decision that put an unqualified person in front of that system and told them the output was their responsibility," he said.
Blame falls on whoever touched the system last
Catherine Bozhenko, Product Manager at DataRobot, said the hunt for someone to blame is itself evidence that nobody designed accountability in the first place. "If your organisation approaches this as a question of who will be blamed afterwards, you have already got it wrong. The reason the blame question comes up at all is usually that the organisation never designed the accountability, so when something goes wrong there is nobody who was formally responsible, and the search begins for whoever touched it last," she said.
She said the person found that way is the easiest to identify and the least useful answer. "That is almost always the engineer, since that is the easiest thing to check, and it is the least useful answer available. An organisation that has designed this properly does not need to ask the question after the event, as the answer was decided before deployment and written down," she said.
Abu Sheikh said responsibility belongs at the top of the organisation. "The responsibility does not sit with the employee. If you are an organisation leader, you should be responsible for whatever your company is exporting, from A to Z, because the employee is part of your organisation and part of your culture. The employee at the end of that chain has inherited whatever decisions were taken above them about how much testing was enough, what the system was allowed to do and what was documented. Holding that person responsible for the consequences of decisions they were never part of is a way of avoiding the question," he said.
Muehmel said where accountability sits depends on how well the system was built. "In some cases it may rest with the end user to make sure they are properly using the technology. In many cases, organisations need to be designing systems that are safe for their employees to use, where the employee essentially cannot use it irresponsibly, and if they are not doing that, then the accountability is on leadership as well," he said.
Accountability belongs on paper before the system goes live
Bozhenko described how organisations that handle this well divide the work. "In most organisations that do these things properly, before you are allowed to put something into production, there is a council of at least four people from four major departments, and then a final person who gives the sign-off on top of that. Each of those four is checking a different category of risk, and none of them is checking work they themselves produced," she said.
She said each seat covers a specific kind of failure. "Usually it is compliance, someone who represents regulation, who checks the documentation, that you evaluated the system properly and that all the problems found are recorded. Then someone from IT, who checks security, authentication, authorisation and the way the system was deployed. Then someone from the business, who ensures the use case makes sense and that the metrics the system was evaluated against make sense for what the business is actually trying to do. And someone from the AI engineers, who double-checks the guards, the tools, the MCP connections and the model that was used. The final sign-off usually comes from the department leader, who takes those four criteria and signs against them," she said.
The principle underneath, she said, is older than AI. "There is a rule that you cannot check your own system and you cannot check your own pull requests. It is never one engineer. The engineer is one of the four voices, checking the part they are qualified to check, and the organisation carries the decision collectively because the decision was collective," she said. She added that none of this works without design work done earlier. "Compliance cannot check documentation that does not exist. The four-person structure only functions where the design work was done, which is why the accountability question and the design question turn out to be the same question asked at two different points in time."
Bozhenko said the same planning decides whether an error becomes a disaster. "It can be a kill switch. You switch off that particular agent, and it is killed, without affecting people and without causing further damage. That capability is the thing that separates an incident from a disaster, and it is a design decision like all of the others. It exists because somebody built it in before it was needed," she said. For a reviewer, the difference is stark. A person with a stop button can halt a failure. A person without one can only watch it and later answer for it.
Muehmel said the assignment of responsibility has to be concrete enough to enforce. "Who is accountable for what may vary across use cases and it may vary across stages, but it absolutely needs to be assigned. It needs to be named and it needs to be tracked," he said. He said the failures themselves also have to travel upwards. "Even if organisations are not sharing publicly that there are failures, they absolutely need to be sharing those failures internally so that there can be shared learnings from them, so that you do not repeat those same failures over and over again within the organisation."
Leaders who skip testing have handed quality control to the most junior person in the room
Abu Sheikh said the executives who approve a system should be the first to test it. "Being the person building the product and being responsible for the company means being with the engineers, in the war zone, building the product with them, testing it with them and validating it with them, instead of only looking down from up high and leaving it at that," he said.
He described the work that should sit above any reviewer. "Before anything reaches a customer, the product or the engine is tested rigorously, with thousands and thousands of simulations run through it. When something comes out in a wrong or harmful way, it is flagged as wrong information, and those responses are exposed only to beta users we have chosen for that purpose, so that the system is being tested with an actual human. All of that work exists above the employee," he said.
Nageswari said her company's leadership speaks directly with the people doing the reviewing. "Our AI leadership team takes a proactive approach when engaging with reviewers working on these AI workflows. This helps them and the product teams understand how outputs are approved or rejected, and the challenges that follow. The aim is for reviewers to approve or reject AI outputs based on context, authority and feedback," she said. Context and authority are the two things the reviewer's role most often lacks.
Abu Sheikh said the fairest version of the job gives the person both the duty and the means to carry it out. "Not everybody is working on responsible AI or ethical AI that will produce a response which is 100% accurate. It is about how you work with it, to ensure that the response coming out is bound by you and not by whatever the AI responded with. An employee who understands that has both the responsibility and the means to discharge it, which is the combination that makes accountability fair," he said.
For the 815 lawyers in Charlotin's database, and for the far larger number of analysts, moderators and claims handlers whose mistakes never reach a court, the second half of that combination is usually missing. They were given the responsibility. Whether the organisations that employ them will also give them the expertise, the visibility and the authority to say no will decide whether this new job becomes a profession or stays a place to put the blame.