Google's ScientistTwo was graded by two AI reviewers. Only one of them was held out, and that is the whole result.
Under the reviewer inside its own refinement loop, ScientistTwo clears the human bar. Under the held-out one, its lead depends on which papers you count.
Contents
A video making the rounds this week asks whether the AI singularity starts with ScientistTwo, a new Google research system. Its title ends in a question mark, which is more care than most coverage of this paper has shown. The system is real, its numbers hold up, and the paper is honest about its own limits. What gets lost between the paper and the summary is which reviewer produced which number.
#What ScientistTwo actually is
ScientistTwo is a real arXiv preprint, "ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI," submitted September 17, 2026. Six of its seven authors, per the paper's own author-affiliation block, list Google Cloud AI Research; the seventh, Yubo Wang, lists the University of Waterloo. It is a fully autonomous multi-agent framework: given a research problem it establishes baselines, forms hypotheses, runs experiments, refines its methods through automated ablation, writes a manuscript, and validates the result through a closed-loop simulated peer-review rebuttal cycle, with no human in the loop mid-run.
The headline numbers hold up under direct check against the paper itself. Across 107 research problems drawn from papers accepted at ICLR 2026, NeurIPS 2025, and ICML 2026, ScientistTwo improved on the human state-of-the-art result in 86 of them, an 80.4 percent success rate, with an average relative improvement of 25.2 percent. Worth pairing with that average: the paper's Table 4 puts the median gain at 7.7 percent, so a small number of large wins are carrying the mean.
#Two reviewers, and only one of them was held out
The paper grades its output with two different automated reviewers, and the distinction between them is the entire story. ScholarPeer is a pre-existing multi-agent peer-review system that ScientistTwo also uses internally, inside its own rebuttal loop, to decide how to revise a draft. The Stanford Agentic Reviewer is a separate pre-existing tool, built by the Stanford Machine Learning Group, used only to grade the finished result. Worth noting how close the in-distribution reviewer sits: ScholarPeer is cited to Goyal et al. 2026, whose author list includes T. Pfister and J. Yoon, both also authors of ScientistTwo. Pre-existing, and from the same lab.
The paper states the asymmetry itself, plainly: ScholarPeer "serves as an in-distribution evaluation, as it is also used to refine the draft quality generated by ScientistTwo. Conversely, Stanford Agentic Reviewer serves as a held-out evaluator that was unseen during development by both the baselines and our method."
Now hold that against the paper's Table 3, which sets ScientistTwo's papers beside real accepted human papers under both reviewers.
Under ScholarPeer, the reviewer inside its own loop, ScientistTwo posts 91.9 percent acceptance against 73.8 percent for accepted human papers. Read the denominators before you read the gap. Table 3's own caption says the "# Papers" column means one thing on the human rows and another on its own: "the number of accepted reference papers evaluated or the number of papers successfully generated by ScientistTwo." The human figure covers all 107 accepted papers. ScientistTwo's covers the 86 papers it actually produced, because on 21 of the 107 problems it finished nothing. Score the same 107 problems and count those 21 as non-acceptances, and ScientistTwo lands at 73.8 percent: level with the humans, on the reviewer inside its own loop.
Under the held-out reviewer, the answer depends entirely on which papers you count, and this is the part worth slowing down for. The paper's claim is that its manuscripts "surpass the average scores of accepted papers at ICLR 2026 and NeurIPS 2025 under both ScholarPeer and Stanford Agentic Reviewer." That claim is true. Under the held-out reviewer ScientistTwo averages 5.4 against humans' 5.2 at ICLR 2026, and 5.6 against 5.5 at NeurIPS 2025. It is also narrow, and by the standard applied further down this page it deserves sizing: the ICLR comparison is five human papers against four of ScientistTwo's, and the NeurIPS margin is a tenth of a point against standard deviations of seven tenths on both sides. At NeurIPS the held-out reviewer also accepts slightly fewer of ScientistTwo's papers than the humans', 75.8 percent against 76.3.
Those two venues are 43 of the 107 problems. The other 64, the largest single group, are ICML 2026 Spotlight papers, and there the held-out reviewer rates accepted human papers 6.1 against ScientistTwo's 5.7, and accepts 62 of the 64 human papers against 34 of the 49 ScientistTwo produced, 96.9 percent against 69.4.
Across the whole set the held-out reviewer accepts 87.9 percent of the 107 human papers and 72.1 percent of the 86 ScientistTwo produced. Put both on the same 107 problems and it is 87.9 percent against 57.9 percent. Those last two conversions are our arithmetic on the paper's own cells, not figures the paper states, and they are exact.
None of this is concealed. The sentence carrying the scoped claim continues in the same breath: "While ScientistTwo does not yet achieve spotlight-level quality, these results still demonstrate that it functions as an expert-level research agent." The project page is careful in the same places, though it never once uses the words "held out." It scopes the same claim to ICLR 2026 and NeurIPS 2025, says outright that "Spotlight-level quality remains out of reach," and frames the 72.1 percent as ScientistTwo being the only agent to clear the held-out reviewer at all, while every other AI system in the comparison scores zero. That last comparison is against other agents, not against people, and it is true.
What travels downstream is the superlative without the scope.
#The per-round view, at its actual size
Table 5 is an ablation on the 49 ICML 2026 Spotlight problems, and it has three rows worth reading together rather than two. With no rebuttal agent at all, ScholarPeer accepts 46.9 percent and the held-out reviewer 49.0 percent. After round one: 79.6 and 73.5. After round two: 93.9 and 69.4.
Read whole, the rebuttal loop lifts both reviewers on net, ScholarPeer by about 47 points and the held-out reviewer by about 20. What diverges is the second round specifically, where ScholarPeer gains another seven papers out of 49 while the held-out reviewer gives back two and its average rating slips from 5.8 to 5.7, inside its own plus-or-minus 0.6 standard deviation. The paper gestures at this without stating it, noting the rebuttal process "improves acceptance rate from Stanford Agentic Reviewer, with this effect being particularly pronounced in the first review round."
Two papers out of 49 is a thin margin and deserves naming at that size. It is a directional hint, not a collapse, and it is the small echo of the much larger venue gap already sitting in Table 3.
#The integrity audit, and who produced the clean number
The authors also run the CoE Integrity Audit over the 49-paper subset: scores reproduce 49 of 49, no specification violations, method matches released code 49 of 49, and zero of 1,814 references hallucinated. That last figure is real, and the ablation rows underneath it are the interesting part. Switch the reference-verification refinement agent off and the same pipeline leaves 19 hallucinated references out of 1,817. Strip all three refinement agents and it is 19 out of 1,840, across 50 papers rather than 49. The clean number is what a dedicated scrubbing agent produces, not what the writer produces unaided, which is worth knowing before it is quoted as proof that the system does not invent citations.
One caution on secondary coverage, and we are following our own advice by not linking it. An auto-generated research note filed two days after the paper cited the same Table 5 figures accurately and argued the pattern undermines the headline result more broadly. It disclosed in its own footer that it was "Generated by Auto Research via a headless coding agent," with sources retrieved by live web search and an explicit "please verify before citing." Its author closed it on September 20, 2026 as superseded by a later run, which was several hours before this note went up.
#What an operator should take from this
Every number we checked in the paper holds up, and the authors scoped their claims honestly. We checked Tables 3, 4, 5 and 7 and the author block, not the whole paper, which is the kind of limit this piece is about. The problem is that one table supports both "beats human papers" and "falls short of human papers" depending on which reviewer and which venue you read, and a summary only ever carries one of them.
That gives you a question you can reuse on any vendor, and it is not specific to this paper. Before trusting a claim that an AI system beats a human benchmark, a reviewer, or a rubric, ask two things: was the system optimized against the same judge being used to declare victory, and what population is the claim scoped to? If the answer to the first is yes, you are reading a training score in an evaluation's clothes. If nobody can answer the second, the superlative has already outrun its evidence.
The failure here is not on the authors. They published both numbers, in the same table, named the held-out reviewer as held out, and stated the limit in the same sentence as the win. It is compression by everyone downstream. The gap between a headline and what a paper's own methodology supports is a recurring pattern worth checking for, and the only way to catch it is opening the results table instead of the summary of it.
#Sources
A new era.
Room for you.
Keep reading
- · 4 min
Anthropic's new project coordinator is real. The team-sharing pitch in the video covering it is not, yet.
Anthropic's Sept 17 Claude Code Projects redesign matches its own docs closely. The video calling it a team OS describes sharing that has not shipped.
- · 9 min
A 40-year-old search algorithm beats modern embeddings at agentic research. Only one version of it does.
A Hornet CEO says BM25 is unreasonably effective for agentic search. The paper he cites disagrees, then a newer paper agrees. The condition between them is the real finding.
- · 6 min
Chunking is not dead. The talk that says so smooths one of its own numbers.
An AI21 researcher attacks query-dependent chunk size with a 17-year-old fusion trick. The mechanism holds. A separate 2026 benchmark found the opposite on a different corpus.