Trust, but verify: AI has made writing cheap, and checking optional

AI drafts the report and nobody checks it. Why "a human in the loop" was never a control, and how the writer's job changed without notice.

A report that once took forty hours to write now takes fifteen. AI drafts it, and the drafts are good. Good enough to be dangerous. But what disappeared with those twenty-five hours was not just typing, but an even more critical control: checking.

A report that once took forty hours to write now takes fifteen. AI drafts it, and the drafts are good. Good enough to be dangerous. But what disappeared with those twenty-five hours was not just typing, but an even more critical control: checking.

When you write something yourself, you check it as you go. Not perfectly, and not always. But you cannot write a claim without at least noticing it, and the weak ones tend to fall apart in your hands before anyone else sees them. That checking was never a separate step. It came free with the writing. The machine does not do it, and the document it produces looks exactly like one that has been verified: polished, fluent, and plausible. That is verification theatre.

It has already cost someone real money. In October 2025 Deloitte refunded part of a AU$440k (US$314k) fee to the Australian government. A report it delivered cited academic papers that did not exist and misstated a Federal Court judgment. The firm's business is assurance. It had used a large language model, and nobody had caught what came back. If a firm that sells checking can publish unchecked work under its own name, the belief that better people would not is wishful.

The machine making mistakes is not the problem. Everyone makes mistakes. The problem is that a check which used to happen automatically now happens only if someone decides to do it, and people are bad at that. Automate a task and you do not free the human – the psychologist Lisanne Bainbridge showed this in 1983. You turn them into a supervisor, and supervising is a poor fit for human attention. Doing something keeps you focused. Watching something does not. Forty years of research on automation bias has confirmed it: people accept machine output even when it is wrong, and miss what the machine left out. Training helps with the first. It does not help with the second.

Even someone who wants to check has less to work with. When you wrote the thing yourself, you knew where every fact came from, because you went and got it. Now you are checking claims you never held in your head, and a fabrication that sounds right looks exactly like a fact that is. The reviewer has less to go on than the writer ever did, and it shows up in three ways:

Polish disarms the reader. Every organisation has one: the charming talker who understands very little, yet speaks with unassailable confidence, and fails upwards. Such individuals are proof that we already reward fluency over substance when the fluency is good enough. The machine has made that fluency free. A clean, confident document still gets less scrutiny, not more, because until recently it usually meant someone had done the work. That signal is gone. Old warning signs, such as the hedge, the seam where two arguments were stitched together, the number that does not sit right, went out with it because the machine is built to sound sure of itself.

Checking the machine with the machine. This one looks like diligence, which makes it worse. The reader asks the AI whether it is sure, and takes the answer as a check. The model may well catch its own mistake. It may not. Either way, it is checking its own work, and any auditor will tell you that is not a check at all. Independence is the whole point – the checker cannot be the thing being checked. An independent check has to land on something the machine did not produce: a source, a record, a fact that is true whether or not the model says so.

The reviewer becomes a postbox. Work arrives, goes through the machine, and is forwarded on, barely read. It looks like laziness. Mostly it is what a busy person does when handed a tool that produces plausible output on demand. Leaning harder on whoever signs at the end will not fix it. A signature at the bottom cannot make up for nobody reading the middle.

A signature at the bottom cannot make up for nobody reading the middle.

The usual answer to all this is a human in the loop. That was never a control, but rather an assumption that a control was happening. When researchers tested it, people relied on the machine's answer even when it was wrong. Explaining the answer did not help, and sometimes made things worse. The only thing that reliably made people think was friction: designs that made it harder to click through without engaging. People rated those designs the worst. Checking is dull work, and without friction, given the choice, they skip it. Putting a human beside the machine does not create a check. It decides who gets blamed when there wasn't one.

More review does not fix this. More review is more theatre. What does fix it is smaller and less comfortable: the check has to belong to someone. Either the person using the machine checks the output as they make it, or someone is named to do it afterwards – and belonging only bites if the checking is built to interrupt, not merely requested. The box below shows what that looks like. What cannot work is the assumption that a signature at the end covers it.

The machine did make the work faster. What nobody said is that it also changed the job. The person who used to write the report now reviews one they did not write, and has less to go on than the writer ever did. The title on the door did not change. The paperwork did not change. Only the work did, and it got harder in the one place nobody was looking.

What actually made people check

The argument above is that checking will not happen by choice. A 2021 study tested what does make it happen. Researchers tried "cognitive forcing functions" – small design frictions that interrupt the reflex to accept a machine's answer. Three worked.

Decide first, see the answer second. The person commits to their own view before the machine's output appears, so they cannot simply anchor on it.

Make them ask. The answer is requested, not served – a deliberate act rather than a default.

Impose a pause. A short enforced wait before the recommendation appears breaks the habit of automatic acceptance.

All three cut reliance on the machine's answer more than showing explanations did. All three were rated worst by the people using them, who also thought their own work had got worse, even as their decisions measurably improved. The discomfort is not a side effect. It is the check working.

Comments

All posts loaded No posts found VIEW ALL Read more Reply Cancel reply Delete By Home PAGES POSTS View All RECOMMENDED FOR YOU LABEL ARCHIVE SEARCH ALL POSTS No posts found Home Sunday Monday Tuesday Wednesday Thursday Friday Saturday Sun Mon Tue Wed Thu Fri Sat January February March April May June July August September October November December Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec Just now 1 minute ago $$1$$ minutes ago 1 hour ago $$1$$ hours ago Yesterday $$1$$ days ago $$1$$ weeks ago more than 5 weeks ago Followers Follow THIS CONTENT IS LOCKED STEP 1: Share to a social network STEP 2: Click the link on your social network Copy All Code Select All Code All codes were copied to your clipboard Can not copy the codes / texts, please press [CTRL]+[C] (or CMD+C with Mac) to copy Table of Content