🔍 Read the full analysis: The Human Cost Of Verifying AI’s Cheap Work on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A source analysis describes a growing gap between AI’s ability to produce work and people’s capacity to verify it, citing examples from mathematics, software and contract workflows. The figures point to review as a potential bottleneck, though some software metrics come from companies that sell review tools, and the long-term effects on training and accountability remain uncertain.
An analysis published this week argues that AI is making work faster to produce while leaving the human effort needed to verify it slow and limited, a mismatch it illustrates with mathematical manuscripts, software changes and contract tasks. The gap matters because organizations may be able to generate more work than experts can reliably review, while responsibility for errors still rests with people and institutions.
The analysis says OpenAI generated 722 mathematical manuscripts from about 4,000 problems, with an average result taking roughly three hours of compute. The papers fall into 372 families. Some results have been formally checked using Lean, a proof-assistant system; OpenAI cautioned that some results without formal verification “could have issues.” The source contrasts that output with the careful review by five leading mathematicians of an earlier result from the programme: a counterexample to an old Erdős conjecture. The example does not establish that all 722 manuscripts have been checked to the same standard.
In software, the analysis cites several industry datasets. Faros AI reported that teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB, using data from 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes took 4.6 times longer to reach the start of review and were accepted at a rate of 32.7%, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with no review during high-adoption periods.
Those numbers have limits: the analysis notes that several cited software-data providers sell code-review tools, and the figures describe different datasets and measures. In contract work, it cites an OpenAI partnership with contract-software company Ironclad. The source says GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks, an improvement over the previous model. That result leaves room for missed requirements, but the source does not provide the full evaluation methodology or task-by-task scores.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If AI output expands faster than review capacity, organizations face a practical limit on how much of that output they can safely use. The bottleneck may shift from drafting code, proofs or contracts to deciding whether each result is correct, relevant and suitable for its intended use. That could raise demand for senior engineers, specialist lawyers, auditors and scientific reviewers, whose approval carries professional or institutional responsibility.
The analysis describes three risks when reviewers cannot keep pace: work may be merged with little or no scrutiny; reviewers may delay AI-generated work because they distrust it; or the producer may effectively decide which results deserve attention. Each response can create costs—missed defects, delays, or dependence on a producer’s own selection process. The cited metrics indicate possible pressure points, not proof that every organization is experiencing them.
There is also a workforce question. The source argues that junior employees often develop judgment by doing the drafting and coding that AI can now help automate. If entry-level workers get less practice producing work themselves, they may have fewer opportunities to acquire the expertise later needed to review it. Whether that effect will occur at scale is not established by the cited data, but it makes training and supervised responsibility relevant to how organizations adopt these tools.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Tension
The examples span different kinds of verification. In mathematics, formal systems such as Lean can check whether a proof follows from stated definitions and assumptions. They do not by themselves determine whether the theorem answers a useful question or whether its framing captures what researchers intended. OpenAI’s warning about unformalized results underscores that the manuscripts do not all have the same level of checking.
Software tests and code review have a similar boundary: tests assess the cases they cover, while reviewers judge whether code fits the broader requirements and system. Contract review adds legal and organizational stakes, including jurisdiction-specific terms and required approvals. The source’s phrase “verification abundance, adjudication scarcity” captures the distinction: automated checks can help, but people still have to decide what the output means and whether it can be relied upon.
The source also points to a disputed mathematical counterexample from August as an example of how a result can require scrutiny beyond checking its presentation. It does not give enough detail here to establish the dispute’s outcome. Across the examples, the common issue is not that machines cannot assist with checking, but that checking a result is different from confirming the question, assumptions and consequences behind it.
software review automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Metrics Cannot Settle
The cited figures do not establish a single, comparable measure of review quality across mathematics, software and contracts. The software results come from different datasets, and some providers have commercial interests in code-review products. The source does not provide full methods for every statistic, so readers cannot determine from these figures alone how much of the change is caused by AI adoption or how representative the organizations are.
It is also unclear how many of OpenAI’s 722 manuscripts were formally verified, independently reviewed or later corrected, and the Ironclad evaluation’s full criteria and task-level results are not supplied. The analysis raises concerns about junior-worker training, but offers no longitudinal evidence showing whether AI use has already reduced the future supply of expert reviewers. Nor does it quantify how much human review can be replaced by improved automated tools.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Track Review and Training
The next useful evidence will be transparent, independent reporting on how AI-generated work performs after review: error rates, time spent checking, rates of correction and the consequences of work that was approved or missed. For software, that means separating review wait times from review quality and comparing teams over consistent periods. For research and contract work, it means publishing enough detail about formal checks, evaluation criteria and later corrections to assess what the headline scores represent.
Organizations adopting AI will also need to decide who is accountable for approving its output and how junior staff gain supervised practice. The source does not identify a policy change or a scheduled follow-up from OpenAI, Faros, LinearB or Ironclad. For now, the reported pattern is an argument to measure verification capacity alongside generation—not evidence that human review has already become an unavoidable limit in every field.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source analysis argues that AI can produce mathematical, software and contract work faster than people can verify it, creating a potential review bottleneck.
Were all 722 mathematical manuscripts formally verified?
No. The source says some results were checked in Lean and quotes OpenAI warning that some unformalized results “could have issues.” It does not state how many manuscripts received formal verification.
What did the software data report?
The analysis cites reports of longer review waits and lower acceptance rates for AI-generated changes in some datasets, as well as a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before being merged or closed. The datasets and measures differ, and some providers sell review tools.
Does this show AI cannot check its own work?
No. The analysis says automated checking can help, but argues that it may not establish whether a result answers the right question or meets broader requirements. It also notes that people and institutions remain responsible for many approvals.
What remains unknown about the effect on jobs?
The source raises the possibility that less drafting work could reduce opportunities for junior workers to build judgment, but it does not provide long-term evidence that this is happening or quantify future demand for reviewers.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
