The Human Cost Of Verifying AI’s Cheap Work
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Human Cost Of Verifying AI’s Cheap Work on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source analysis describes a growing gap between AI’s ability to produce work and people’s capacity to verify it, citing examples from mathematics, software and contract workflows. The figures point to review as a potential bottleneck, though some software metrics come from companies that sell review tools, and the long-term effects on training and accountability remain uncertain.

An analysis published this week argues that AI is making work faster to produce while leaving the human effort needed to verify it slow and limited, a mismatch it illustrates with mathematical manuscripts, software changes and contract tasks. The gap matters because organizations may be able to generate more work than experts can reliably review, while responsibility for errors still rests with people and institutions.

The analysis says OpenAI generated 722 mathematical manuscripts from about 4,000 problems, with an average result taking roughly three hours of compute. The papers fall into 372 families. Some results have been formally checked using Lean, a proof-assistant system; OpenAI cautioned that some results without formal verification “could have issues.” The source contrasts that output with the careful review by five leading mathematicians of an earlier result from the programme: a counterexample to an old Erdős conjecture. The example does not establish that all 722 manuscripts have been checked to the same standard.

In software, the analysis cites several industry datasets. Faros AI reported that teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB, using data from 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes took 4.6 times longer to reach the start of review and were accepted at a rate of 32.7%, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with no review during high-adoption periods.

Those numbers have limits: the analysis notes that several cited software-data providers sell code-review tools, and the figures describe different datasets and measures. In contract work, it cites an OpenAI partnership with contract-software company Ironclad. The source says GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks, an improvement over the previous model. That result leaves room for missed requirements, but the source does not provide the full evaluation methodology or task-by-task scores.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA source analysis argues that rapid growth in AI-generated work is outpacing the human capacity to check and take responsibility for it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

If AI output expands faster than review capacity, organizations face a practical limit on how much of that output they can safely use. The bottleneck may shift from drafting code, proofs or contracts to deciding whether each result is correct, relevant and suitable for its intended use. That could raise demand for senior engineers, specialist lawyers, auditors and scientific reviewers, whose approval carries professional or institutional responsibility.

The analysis describes three risks when reviewers cannot keep pace: work may be merged with little or no scrutiny; reviewers may delay AI-generated work because they distrust it; or the producer may effectively decide which results deserve attention. Each response can create costs—missed defects, delays, or dependence on a producer’s own selection process. The cited metrics indicate possible pressure points, not proof that every organization is experiencing them.

There is also a workforce question. The source argues that junior employees often develop judgment by doing the drafting and coding that AI can now help automate. If entry-level workers get less practice producing work themselves, they may have fewer opportunities to acquire the expertise later needed to review it. Whether that effect will occur at scale is not established by the cited data, but it makes training and supervised responsibility relevant to how organizations adopt these tools.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Tension

The examples span different kinds of verification. In mathematics, formal systems such as Lean can check whether a proof follows from stated definitions and assumptions. They do not by themselves determine whether the theorem answers a useful question or whether its framing captures what researchers intended. OpenAI’s warning about unformalized results underscores that the manuscripts do not all have the same level of checking.

Software tests and code review have a similar boundary: tests assess the cases they cover, while reviewers judge whether code fits the broader requirements and system. Contract review adds legal and organizational stakes, including jurisdiction-specific terms and required approvals. The source’s phrase “verification abundance, adjudication scarcity” captures the distinction: automated checks can help, but people still have to decide what the output means and whether it can be relied upon.

The source also points to a disputed mathematical counterexample from August as an example of how a result can require scrutiny beyond checking its presentation. It does not give enough detail here to establish the dispute’s outcome. Across the examples, the common issue is not that machines cannot assist with checking, but that checking a result is different from confirming the question, assumptions and consequences behind it.

Amazon

software review automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Metrics Cannot Settle

The cited figures do not establish a single, comparable measure of review quality across mathematics, software and contracts. The software results come from different datasets, and some providers have commercial interests in code-review products. The source does not provide full methods for every statistic, so readers cannot determine from these figures alone how much of the change is caused by AI adoption or how representative the organizations are.

It is also unclear how many of OpenAI’s 722 manuscripts were formally verified, independently reviewed or later corrected, and the Ironclad evaluation’s full criteria and task-level results are not supplied. The analysis raises concerns about junior-worker training, but offers no longitudinal evidence showing whether AI use has already reduced the future supply of expert reviewers. Nor does it quantify how much human review can be replaced by improved automated tools.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review and Training

The next useful evidence will be transparent, independent reporting on how AI-generated work performs after review: error rates, time spent checking, rates of correction and the consequences of work that was approved or missed. For software, that means separating review wait times from review quality and comparing teams over consistent periods. For research and contract work, it means publishing enough detail about formal checks, evaluation criteria and later corrections to assess what the headline scores represent.

Organizations adopting AI will also need to decide who is accountable for approving its output and how junior staff gain supervised practice. The source does not identify a policy change or a scheduled follow-up from OpenAI, Faros, LinearB or Ironclad. For now, the reported pattern is an argument to measure verification capacity alongside generation—not evidence that human review has already become an unavoidable limit in every field.

Amazon

contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The source analysis argues that AI can produce mathematical, software and contract work faster than people can verify it, creating a potential review bottleneck.

Were all 722 mathematical manuscripts formally verified?

No. The source says some results were checked in Lean and quotes OpenAI warning that some unformalized results “could have issues.” It does not state how many manuscripts received formal verification.

What did the software data report?

The analysis cites reports of longer review waits and lower acceptance rates for AI-generated changes in some datasets, as well as a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before being merged or closed. The datasets and measures differ, and some providers sell review tools.

Does this show AI cannot check its own work?

No. The analysis says automated checking can help, but argues that it may not establish whether a result answers the right question or meets broader requirements. It also notes that people and institutions remain responsible for many approvals.

What remains unknown about the effect on jobs?

The source raises the possibility that less drafting work could reduce opportunities for junior workers to build judgment, but it does not provide long-term evidence that this is happening or quantify future demand for reviewers.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Danganronpa 2X2 – Official Release Date Trailer | Nintendo Direct September 2026

The official trailer for Danganronpa 2×2 aired during Nintendo Direct September 2026, confirming the game’s release date and fueling fan anticipation.

BeamNG.drive Climbing The Steam Charts

BeamNG.drive has climbed to the third most-played game on Steam, reaching a peak of over 24,000 players, marking a significant surge in its popularity.

AI’s Inner Engine: Inside Twelve Machines Driving Innovation

An in-depth look at twelve key AI machines driving technological progress, explaining how they work, their significance, and what remains to be understood.

How Tech Giants Are Pioneering AI Development

An analysis of how major tech companies are pioneering AI, facing platform shifts that could redefine dominance and survival in the industry.