Claude Leads a Quarter of Anthropic’s AI Research: What the Number Is Worth

Bildschirm mit Programmcode in einem Forschungslabor als Sinnbild für automatisierte KI-Entwicklung
Photo by Kevin Ku on Unsplash

On September 17, 2026, Anthropic published its first figures on how much of its own AI development is now carried by Claude itself. According to the company, the model “leads” 26 percent of its research and development work, up from under one percent in February. The number is striking – but it comes from the vendor, was produced with the help of its own models, and measures something different from what the headline suggests.

Key takeaways

  • Anthropic uses an “R&D Automation Index” to measure how much of its AI research runs at which level of automation; at 26 percent (as of August 2026), Claude does most of a task from start to finish while a human supervises.
  • For more than 90 percent, the AI works at least “in collaboration” with humans; according to Anthropic, no measured area runs fully autonomously.
  • The 26 percent refers to a basket of tasks weighted by working time, not to 26 percent of all individual tasks.
  • A Claude model acts as the judge in the ratings; Anthropic itself concedes that independent verification by third parties would make sense.
  • Anthropic also reports figures on oversight of AI agents and on the computing power devoted to safety work.

What the index measures

The basis is a scale developed by the research institute Epoch AI, running from level AL0 (“no AI involvement”) to AL5 (“fully autonomous, no human in the loop”). At AL3 (“collaborates”), the AI does large chunks of the work under close human direction. At AL4 (“leads”), it solves most of a task on its own from a brief prompt, and a human reviews. Anthropic gives an example: if a nightly data pipeline breaks, Claude at this level finds the cause, writes and tests the fix, and documents anything surprising. The engineer only reads the write-up and decides whether to ship. Claude does not deploy – that would be AL5.

According to Anthropic’s publication, Claude does not reach AL5 in any measured area. The 26 percent at AL4 is still a leap: in February 2026 the figure was under one percent, according to the chart caption. Across all levels from AL3 up, Anthropic arrives at more than 90 percent of the work examined.

How the number is produced – and where it wobbles

The method explains why the figure deserves caution. For each week in July 2026, Anthropic randomly drew 20 percent of the staff in each department involved in model training. A Claude agent read through their week in Slack and internal documents and listed their activities – about 15,000 individual tasks. Claude sorted them into a tree of 542 nodes, 378 of them leaves such as “evaluation platform defect diagnosis” or “network policy for RL sandboxes.” A second Claude agent researched how the work actually gets done for each node, and a separate Claude model acting as judge assigned the automation level. Weighting follows working time: nodes where humans spend more time count for more. As heise online stresses, the 26 percent is therefore a share of this weighted basket and not a quarter of all recorded tasks.

Anthropic names the weaknesses itself. When compared with people who own the relevant areas, the judge model matched them exactly in 59 percent of cases; humans agreed with each other in only 35 percent. In 97 percent of cases the ratings were at most one level apart. The line between “collaborates” and “leads” thus remains a matter of interpretation. There is also the circularity problem: a Claude model rates Claude’s work. The task basket is also frozen, so the index cannot show whether humans are taking on new kinds of work. A test against a comparison basket from January found no rise in new tasks, according to Anthropic. As a remedy, the company proposes checks by third parties or by models from other developers.

Oversight and safety budget

The publication contains two further measures. On oversight of AI agents, Anthropic cites about 30,000 agents doing research and engineering at the same time on its most-used internal platform. A real-time monitor reviews every action before execution, according to the report; of more than a billion decisions in August, roughly one in 47,000 was blocked (0.002 percent). Offline monitors flag about 100,000 transcripts a week, and about 50 of them with the highest priority go to humans. The figures apply to this one platform only and come from Anthropic itself.

On computing power, in a sample week in July about 6 percent of the capacity used for AI research went to safety tasks, and about 12 percent for AI-driven AI research. Anthropic calls the estimate deliberately conservative and points out that a single week shows no trend and that the compute share does not measure how effective the safety work is.

Assessment and outlook

What is being measured here is a precursor to a question the industry is wrestling with: how far is a lab from an AI that develops its successors largely on its own (“recursive self-improvement”)? Anthropic wants to publish the metrics regularly and admit independent evaluators into the company; the figures are expected to shift if the industry slows down as Dario Amodei has demanded. OpenAI, too, reported an automated research intern in early September while warning of alignment risks; how closely safety work and automation are linked was shown earlier by Anthropic’s experiment on aligning other models.

The real value for now lies not in the size of the number but in the fact that a lab discloses a traceable method at all. As long as only one vendor measures with its own models, however, it remains open whether 26 percent is a lot or a little. The index becomes meaningful only when other labs replicate it and third parties check the results. Until then it is a self-report with a transparency ambition – a start, but not yet proof.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top