Sunday, 16 August 2026
The Verified Journalism Press

Journalism with its sources attached.

Sections
WORLD
AUSTRALIA
INDIA
BUSINESS
TECHNOLOGY
SCIENCE
SOCIETY
RIGHTS
CORRUPTION
CULTURE
OPINION
FAMOUS
The Press
Latest
Brussels has child safety cases open against Snapchat, Meta and TikTok, but not YouTube or the app storesMost Australian under-16s are still using social media, the regulator's own evaluation findsAI-designed viruses clear peer review, then an independent check finds them close relatives of the natural originalMIT's AI supercomputer has fallen 36 places in the world rankings without getting any slowerArizona physicists shift the quantum noise inside a light pulse, and watch it move in real timeApple has handed Siri to Google, and Amazon's Alexa+ has reached AustraliaBrussels has child safety cases open against Snapchat, Meta and TikTok, but not YouTube or the app storesMost Australian under-16s are still using social media, the regulator's own evaluation findsAI-designed viruses clear peer review, then an independent check finds them close relatives of the natural originalMIT's AI supercomputer has fallen 36 places in the world rankings without getting any slowerArizona physicists shift the quantum noise inside a light pulse, and watch it move in real timeApple has handed Siri to Google, and Amazon's Alexa+ has reached Australia
Markets
ASX 200
S&P 500
Nasdaq
FTSE 100
Nikkei
Gold
Brent
AUD / USD
AUD / EUR
AUD / GBP
AUD / JPY
Bitcoin
Ethereum
Yahoo · ECB · CoinGecko

Front page / Research

AI and research integrity

AI agents re-ran the 168 top papers at a machine learning conference; eight held up above 80 per cent

SAI Labs reported on 22 July 2026 that its agents attempted to replicate every oral paper at ICML 2026, completing 105 full reruns. Of the 92 papers with at least five checkable claims, 34 reproduced more than 40 per cent of attempted claims and eight reproduced more than 80 per cent.

Stanford University, William Gates Computer Science Building, May 28, 2023 - 75
Stanford University, William Gates Computer Science Building, May 28, 2023 - 75. Photograph: 94rain, CC BY-SA 4.0

SAI Labs, a for profit research review company registered in Delaware, published an analysis on 22 July 2026 in which artificial intelligence agents attempted to verify all 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning. The agents extracted each paper's central claims, downloaded the authors' code and data, re-ran experiments where the resources allowed, and compared the output with what the paper reported. One hundred and four of the papers shipped runnable code. One hundred and five full replications were completed.

The headline result is narrow and should be read carefully. Of the 92 papers with at least five claims the agents could assess, 34 reproduced more than 40 per cent of the claims that were attempted, and eight reproduced more than 80 per cent. Nature, reporting the analysis on 6 August, described the same figure as the agents being able to reproduce more than two of five claims from only 34 papers. Those are the same underlying numbers expressed differently, and neither is a statement that the remaining papers are wrong. SAI's own report distinguishes claims that were not attempted from claims that ran at reduced scope and claims that ran in full, and the company says its review system "is still in beta and it can make mistakes".

The cost figures are the part that has attracted least attention and may matter most. SAI put the median cost of replicating a paper at about USD 8,900. Seventeen papers cost more than USD 100,000 to reproduce, and the most expensive came to roughly USD 2.2 million in compute. Those estimates exclude retries, salaries and infrastructure, so they are conservative by the company's own account. If verification is that expensive, no conference programme committee staffed by volunteers was ever going to do it, which is the argument SAI is in business to make.

SAI also compared its agents against the conference's human reviewers. It reported that 78 per cent of the comments two or more human reviewers agreed on were also covered by its review, that its agents raised 903 code and reproducibility issues the human reviewers did not, and that human reviewers found 22 issues the agents missed. That comparison is drawn by the vendor, from its own logs, and has not been independently audited.

A separate line of work points the same way. In a preprint posted on 5 December 2025, Federico Bianchi, Yongchan Kwon, Zachary Izzo, Linjun Zhang and James Zou of Stanford University and collaborators built a checker on GPT-5 and ran it over published machine learning papers. They found objective errors, in formulae, calculations, tables and internal contradictions, rising from 3.8 per paper at NeurIPS in 2021 to 5.9 in 2025, an increase of 55 per cent. Human reviewers confirmed 83.2 per cent of the issues the tool flagged. The authors deliberately excluded judgements about novelty and significance. "AI should not do everything, and leave choices about novelty and significance to humans," Bianchi told Nature.

The case against treating these tools as arbiters is also on the record, and it comes from a benchmark. SPOT, released in May 2025 by Guijin Son and colleagues, pairs 83 published papers with 91 errors serious enough to have triggered an erratum or a retraction. The best model tested reached 21.1 per cent recall at 6.1 per cent precision, most performed close to zero, and across eight independent runs the models rarely rediscovered the same errors. Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology, told Nature the tools "make mistakes like humans do", and that their output needs human oversight.

Progress against that benchmark is being claimed. On 26 June 2026 a Google team led by Rajesh Jayaram described a Paper Assistant Tool that checks theoretical results and validates experiments, reporting a 34 per cent improvement over zero shot recall on mathematical errors in SPOT, and pilot deployments as a pre-submission review tool at the STOC and ICML conferences. Separately, on 10 August 2026, Paul Litvak planted 100 errors across 62 categories into ten open access psychology papers and tested detection systems against them. The best single model caught 71, the worst 30, and an ensemble of foundation models caught 91. Litvak measured recall only, not false positives, which he says is the study's main limitation.

What none of this settles is the question the numbers appear to answer. Failure to reproduce a claim inside an agent's compute budget is not proof the claim is false, and the organisation reporting the failure rate for machine learning research sells machine learning research review. SAI has not published a peer reviewed account of its method, ICML has not commented on the findings, and no independent group has re-run the audit of the auditors.

Sources

Every factual claim above rests on the 8 published sources below. They are listed so you can check the reporting rather than take it on trust.

  1. SAI LabsHow much science is verifiable? We replicated ICML 2026 oral papers
  2. NatureAI agents are checking the scientific literature and spotting decades-old errors
  3. arXivTo Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
  4. arXivWhen AI Co-Scientists Fail: SPOT, a Benchmark for Automated Verification of Scientific Research
  5. arXivTowards Automating Scientific Review with Google's Paper Assistant Tool
  6. Paul LitvakHow well does AI peer review work?
  7. ScienceHundreds of paper-mill papers peddled in ads were later published
  8. Retraction WatchWeekend reads: Jason Arday; research gold; AI; OpenAI math breakthroughs

The Verified Briefing

One email each morning. Every story in it carries its sources, so you can check the reporting before you repeat it.

No tracking pixels. One click to leave.