Back to Blog
AI & Automation
By Anil Konur
July 28, 2026

AI Made the Draft Cheap. It Made the Check Expensive.

Bad AI output gets caught, because it looks bad. Plausible AI output does not - and a review gate that only works when your best person is in the chair is not a gate.

In April I wrote AI and the White-Collar Workforce: From Threat to Catalyst, and one line in it has held up better than the rest: AI does not admit uncertainty. It always offers an answer, accurate or not.

I stand by that. What I did not do was say what to build because of it. This post is that.

The standard version of this argument, the one you have read a hundred times by now, is garbage in, garbage out. Your AI output is only as good as your inputs, so hire good people and write better prompts. It is a comforting claim and it is going stale. Models now produce competent work from lazy prompts. The input quality problem is real, and it is shrinking.

The problem that is not shrinking is a different one. Plausible in, plausible out.

The Konur Consulting take: AI collapsed the cost of producing work. It did not collapse the cost of knowing the work is right. The bottleneck moved from production to verification, and most operations are still staffed, measured, and budgeted for the old bottleneck.

The failure mode is not wrong. It is plausible.

Bad output is not the risk. Bad output gets caught, because it is obviously bad.

The risk is output that is correctly formatted, confidently worded, structurally sound, and wrong in a way that only surfaces downstream. The brief citing a case that does not say what it is claimed to say. The application that demos beautifully and has no authorization model behind it. The analysis with a defensible-looking number resting on a definition nobody checked.

Polish used to carry information. Someone who took the time to format it properly had probably taken the time to get it right. AI severed that correlation. Polish now costs nothing, so it tells you nothing. But we all still read it as a signal, which is precisely why fluent wrong output moves through an organization faster than clumsy wrong output ever did.

Why "have a smart person check it" is not a plan

The obvious response is to put an experienced person in front of the output. It is the right instinct and an incomplete answer, for three reasons.

It does not scale. One person can carefully review five documents. They cannot carefully review four hundred. When volume arrives, careful review quietly becomes skimming, and nobody announces the switch. The process on paper is unchanged. The process in practice has stopped working.

Experience can cut the wrong way. A senior reviewer often approves AI output faster, not slower, because it reads like what they would have written themselves. Recognition feels like verification. It is not. The reviewer best equipped to catch a subtle error is also the one most likely to be disarmed by a familiar-sounding one.

It leaves when the person leaves. Quality that depends on a specific individual's judgment is quality you lose to a resignation, a vacation, or a bad week.

This is a familiar shape. Capability that lives in a person is heroics. Capability that lives in a system is infrastructure. That distinction has been the center of how I think about scaling operations for twenty years, and AI did not change it. It raised the stakes.

What a verification gate actually looks like

A gate is not a meeting or a culture of rigor. It is four specific things, and they fit on one page.

Acceptance criteria written before generation. Decide what correct means for this output before you see the output. Judging quality afterward, against a draft that is already busy persuading you, is not a test. It is a negotiation you will lose.

A named owner for verification. Not "the team reviews it." A person, on the record, for each class of output. Diffuse responsibility for checking is functionally the same as no responsibility for checking.

A provenance requirement. Every factual claim, citation, figure, and dependency traces to a source someone can open. This single rule catches most of what actually goes wrong, and it works whether the reviewer is brilliant or ordinary. That is the entire point of writing it down.

Sampling and audit on volume work. Where output is high-volume, nobody reviews all of it. Accept that rather than pretending otherwise: sample a defined percentage, log what the sample finds, and route the failure patterns back into the acceptance criteria. Unaudited volume is not efficiency. It is unpriced risk.

Notice what none of these require. None of them require a genius in the chair. That is the design goal. A gate that only works when your best person is running it is not a gate.

The honest part

The set of tasks that genuinely require an expert reviewer is shrinking. Pretending otherwise is its own form of hype, and the "AI still needs humans" genre has quietly become a comfort product for people who feel exposed by the technology.

So the useful question is not whether AI still needs people. It is narrower: which decisions do we still gate, why those specifically, and how would we know if we had gated the wrong ones?

Answering that honestly means some gates come down. An organization that gates everything is not being careful. It has just declined to think about where the risk actually sits, and it will pay for that in speed while getting no safety in return.

What to do Monday

  • Pick one high-volume AI-assisted output and write down what correct means for it, before anyone generates the next one.
  • Name the verifier. One person, per output class, on the record.
  • Require provenance on claims. Every figure, citation, and dependency traces to something openable. Make it a submission requirement, not a review preference.
  • Set a sampling rate for anything you cannot review completely, and log what the sample turns up.
  • Audit your gates for cost, not just presence. Find one review step that has never caught anything and remove it. A gate that never fails anything is theater, and it teaches people that gates are theater.

FAQ

Isn't this just garbage in, garbage out?

No, and the difference is the whole argument. Garbage in, garbage out says weak inputs produce weak outputs, which is true and increasingly easy to avoid. Plausible in, plausible out says good-looking inputs produce good-looking outputs that may still be wrong. The first is an input problem you can fix with better prompts. The second is a verification problem, and you can only fix it with process.

Doesn't a review gate cancel the speed gain from AI?

It costs some of it, and that is the correct trade. The alternative is not faster. It is faster until the first expensive failure and considerably slower after. The gate is also cheaper than it sounds, because most of the work is defining acceptance criteria once, not reviewing every artifact forever.

We are a small team. Isn't this overhead we cannot afford?

The version described here is one page: acceptance criteria, a named owner, a provenance rule, a sampling rate. Small teams are where the person-dependent alternative fails hardest, because there is no bench when that person is unavailable.

Anyone can produce the artifact now. The advantage belongs to whoever can tell, reliably and repeatably, whether the artifact is right.

Konur Consulting helps organizations operationalize AI so it holds up under scrutiny: acceptance criteria, review gates, provenance discipline, and the audit loops that keep quality from depending on who happens to be in the room. If your AI adoption has produced more output than confidence, that is where to start. Reach out at info@konurconsulting.com to start the conversation.