Verification Is the New Bottleneck in Enterprise AI
When machines produce faster than anyone can review, the constraint moves.
For most of a decade, one public health dataset — America’s NHANES survey — produced a steady trickle of a particular kind of research paper: single-factor association studies, about four a year, every year, from 2014 through 2021. Then generative AI arrived. Thirty-three in 2022. Eighty-two in 2023. One hundred ninety in 2024 — and that count stopped on October 9th, because that’s when the University of Surrey team ran the search.
Production scaled nearly fifty-fold. The world’s supply of qualified reviewers did not move. And the Surrey study found exactly what you’d expect when output outruns review: formulaic manuscripts, designs that ignore how multifactorial health actually is, false discoveries sliding through, journals overwhelmed. Science just got there first — because science counts its output in public. Your company is running the same experiment right now. You simply can’t see the curve.
Every bottleneck we break moves the bottleneck
Seventy years of computing has been one long campaign against production constraints. Compute got cheap, then storage, then distribution, and now — with generative AI — the production of text, code, analysis, and decisions itself. Each time, we celebrated the abundance and then discovered the oldest rule in operations still applied: a system’s throughput is set by its bottleneck, and accelerating any other step doesn’t create output. It creates inventory.
That is precisely what AI has just done to knowledge work. It supercharged the production step, and the unreviewed output is now piling up in front of the one station that didn’t speed up: the person qualified to say this is correct, this is safe, this is ours. Code waiting for review. Contracts waiting for counsel. Analyses waiting for someone who can tell insight from artifact. And agent actions waiting — increasingly — for no one at all, which is its own essay in this series.
Quality is the dimension that didn’t get faster
Google’s engineering productivity research team has measured developers for years on three deliberately counterbalancing dimensions — speed, ease, and quality — because improving one at the expense of the others is trivial. One of the team’s leads liked to make the point with a joke: the fastest way to boost speed and ease is to delete code review. Everyone laughs, because everyone understands that’s not productivity — it’s shipping terrible software faster.
Now look at what AI actually did to that triad. It collapsed the cost of speed. It collapsed the cost of ease. It left quality as the one surviving human-gated dimension — the joke, industrialized. Any organization measuring only velocity right now is watching two dials spin and calling it progress, while the third dial is the only one that still separates them from the NHANES chart.
The organizational consequence is an inversion. Producers used to outnumber verifiers because production was the expensive part. When production approaches free, your best people stop producing and become full-time reviewers — and the apprenticeship model quietly breaks, because juniors used to learn by producing under review, and now the machine produces while the junior watches. “Human in the loop” becomes, in practice, human as the loop: the throughput ceiling of the whole system.
Activity is not an outcome until it’s verified
This is where the measurement conversation every boardroom is having actually lands. The first year of enterprise AI was measured in inputs — seats activated, tokens consumed, usage leaderboards. Boards have since remembered what budgets are and moved to the harder question: what are we getting back? The honest answer is that usage counts activity, and an outcome only exists once something verifies it. Verification is the conversion step between raw output and value — between a draft and a deliverable, between an agent’s action and a result you would put your name on.
In the last essay I argued that outcome pricing demands outcome accountability. Verification is where that accountability gets manufactured. And for agent actions, the problem is sharper than for agent documents: a wrong paragraph waits harmlessly until somebody reads it, but a wrong action has already happened. Verifying actions means evidence captured at the moment of the act — including the authorization decision itself, down to the data element. That evidence layer is the half of this problem I’ve been working on in the open, in a proposed OCSF extension on GitHub: the same record that satisfies an auditor is what lets verification scale past the humans doing it by hand.
What to do now
Five moves, none of which require waiting for better models:
Map your verification stations. Every place where “someone has to check this” lives — code review, legal review, QA, compliance sign-off, financial approval — and measure the queue in front of each. Queue depth at the review station, not license count, is your real AI capacity.
Invest at the bottleneck, not the production step. Every additional dollar spent generating output before the review station scales is a dollar spent manufacturing inventory. Evaluation tooling, review tooling, and verifier time are where the throughput is.
Make quality a first-class metric beside speed and ease. Google’s triad exists because the dimensions discipline each other. If you measure only velocity, AI will happily give you velocity.
Tier verification by risk. Sample the low-stakes output; gate the high-stakes. Write down which classes of output and action require human sign-off, and which can accept evidence-backed automated checks. An undifferentiated review queue fails both ways at once.
Turn review into evidence. A verification that lives in a reviewer’s head evaporates the moment it happens. Recorded verification decisions compound — audits, outcome pricing, and customer trust all draw on the same ledger.
And treat the five as a loop, not a checklist: evidence that never flows back into a decision — renewing an authority, restricting it, suspending it — is a report, not assurance. The loop closes when what verification finds changes what the system is allowed to do next, with a named business owner accountable for its continued use.
The verification dividend
Constraint economics has a generous ending for whoever industrializes the constraint. The company that can verify at scale gets to safely use more AI than everyone else — more delegation, more automation, more outcome-priced deals it can actually sign, because it can prove what its systems did. That’s the same conclusion this series keeps arriving at from different directions, and it isn’t a coincidence.
Production is becoming table stakes. Proof is becoming the product.
Sources
- Suchak, T., Aliu, A.E., Harrison, C., Zwiggelaar, R., Geifman, N., Spick, M. — Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health database, PLOS Biology 23(5), 2025: journals.plos.org
- Nature (news) — AI linked to explosion of low-quality biomedical research papers (May 2025): nature.com
- The New Stack — How Google Unlocks and Measures Developer Productivity (Jaspan & Green on the speed / ease / quality triad): thenewstack.io
This is the third essay in Field Notes on the Agentic Enterprise. Previous: Software Sold Seats. Cloud Sold Consumption. AI Will Sell Outcomes. Next: “AI Doesn’t Fix Operational Debt. It Compounds It.”
Helping revenue leaders, founders, and investors build the future of go-to-market.
© 2026 Todd Yancey. All rights reserved.
