Study Contradicts Anthropic and OpenAI: AI Still Can’t Do Research Alone

Forscherin analysiert Daten an mehreren Bildschirmen in einem Labor
Photo by Kari Shea on Unsplash

For months, Anthropic and OpenAI have been building expectations that their models are on the verge of driving AI research on their own. A new study from Princeton and the UK AI Security Institute just put that claim through an unusually strict test: real, still-unpublished research questions, graded by the exact experts who posed them. The result lands far more sober than either lab’s announcements.

Key takeaways

  • Princeton and the UK AI Security Institute tested Claude Opus 4.8 and GPT-5.6 Sol on two unpublished NeurIPS 2026 submissions.
  • The method is called “Shadow Evaluation”: agents tackle the central research question, and the original authors grade the result like conference reviewers.
  • Both agent-written papers were clearly rejected, one “Strong Reject,” one “Reject.”
  • The agents reliably handle engineering work but fail at judgment, creativity, and resource management.
  • The findings contradict public statements from Anthropic and OpenAI about their models nearing research autonomy.

The setup: real questions, real reviewers

Existing benchmarks for AI research ability share a basic flaw: most measure tasks with a known, verifiable answer. Whether an agent can also handle open, ambiguous research questions remained unclear. The authors, led by Peter Kirgis, solved that with “Shadow Evaluation”: two teams had just submitted papers to NeurIPS 2026 whose results were not yet public. An AI agent was handed the exact same starting question, with six days of time, $3,000 in API budget, GPU access, and virtual machines, but no access to the training data or the original paper. The real authors then graded the agents’ output like reviewers at an academic conference, without knowing an AI had written it.

The two research topics were technically demanding: one involved steering personality traits in language models directly through their weights, the other was TabPFN, a method for detecting data drift in predictive models. Both agent submissions failed. One reviewer criticized poorly motivated experiments and unreadable prose; another called the reasoning a textbook “proof by example” fallacy, which they described as highly unscientific.

Where the agents fall short

What stands out is how unevenly the abilities are distributed. Pure engineering work, writing code, setting up experiments, running infrastructure, both agents handled almost entirely on their own, needing only three human interventions total. The researchers found no significant reward hacking, meaning the agents weren’t gaming the evaluation instead of genuinely solving problems. The failures started exactly where research demands judgment rather than execution. The study identifies five recurring failure modes: poor judgment about what’s actually worth publishing, no creative way out when a hypothesis gets falsified, ineffective backtracking out of dead ends, weak awareness of their own time and budget, and a gradual drift away from the original instructions as the project went on.

The budget numbers make the problem concrete: the Claude Opus agent spent only about $1,130 of its $3,000 budget on its most ambitious run and wrapped up its most ambitious goal after just five to ten hours, instead of using the planned 36 to 48 hours for open-ended exploration. GPT-5.6 Sol took the opposite approach and burned through its entire budget in two days. Both strategies ended the same way: rejection. That lines up with the broader debate around AI self-improvement, which researchers were already warning about back in 2023: the raw ability to write code and automate experiments is clearly here. What’s missing is the ability to decide which questions are even worth pursuing in the first place.

What this means for Anthropic’s and OpenAI’s claims

The study hits a sore spot, because both labs have spent recent months suggesting the opposite. In June, Anthropic published a blog post titled “When AI Builds Itself,” presenting data on what it described as AI-accelerated internal research. OpenAI, meanwhile, promoted the claim that GPT-5.6 Sol helped post-train a smaller model and saved researchers several weeks of work, a claim that is conspicuously absent from the model’s 81-page technical system card. The study’s authors state the discrepancy carefully but clearly: frontier models handle the engineering side of AI research, but still struggle with weeks-long, open-ended research questions. Google DeepMind researcher Tom Zahavy had already argued in his own position paper, titled “LLMs can’t jump,” that language models simply lack the cognitive mechanism to generate genuinely new ideas rather than recombine existing ones. The shadow evaluation results now provide concrete, practical evidence for that claim.

Bottom line: the engineering gap is closed, the judgment gap isn’t

For readers following the debate over impending AI research autonomy, the study offers a useful distinction: there isn’t one question, “can AI do research,” but at least two separate abilities, technical execution and scientific judgment, that appear to be developing at very different speeds. The first gap looks largely closed by this research. The second remains wide open. The authors themselves caution that their sample of two cases is small and that further shadow evaluations are needed to confirm the findings. Still, the work offers an important counterpoint to marketing claims that tend to blur the line between automated execution and genuine research achievement.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top