A Small AI Model Claims It Can Out-Research Anthropic and OpenAI

Wissenschaftlerin analysiert Forschungsdaten am Bildschirm im Labor
Photo by Julia Koblitz on Unsplash

A London startup with roots at Google DeepMind claims its comparatively tiny AI model has beaten the much larger systems from Anthropic and OpenAI at a scientific task. Inherent has unveiled its agent Faraday, which independently reproduces published research papers, and says it outperforms Claude Opus 4.8 and GPT-5.5. The catch: so far, every number comes from the company itself.

Key takeaways

  • Inherent was founded by former Google DeepMind researchers led by chief scientist Edward Hughes and is based in London.
  • Its agent Faraday runs on the comparatively small Qwen 3.6 language model with 27 billion parameters, but relies on OpenAI’s GPT-5.5 Codex for the coding portions of its work.
  • On Inherent’s own Replica benchmark of 310 tasks drawn from 100 published papers, Faraday says it won 73 percent of in-distribution test cases and 60 percent of tasks from entirely new, held-out papers against Claude Opus 4.8 and GPT-5.5.
  • Faraday was trained with reinforcement learning, which rewards successful approaches instead of prescribing rigid rules, aiming to give the model something like scientific intuition.
  • No independent verification of the results exists yet; Inherent has not published the specific test papers or its full evaluation methodology.

What Faraday actually does

The task sounds like a rite of passage for early-career scientists: an agent receives a published paper stripped of its figures and result charts, and has to reconstruct them independently from the text description alone. Co-founder Edward Hughes told TechCrunch this is a ‘standard exercise for human scientists,’ comparable to starting a PhD. Inherent’s Replica benchmark covers 310 individual tasks from 100 papers in machine learning, materials science, and weather forecasting, each performed under a limited time and compute budget.

Small beats big, at least on paper

What stands out most is the size gap: Faraday runs on a 27-billion-parameter base model, far smaller than the frontier systems Claude Opus 4.8 and GPT-5.5 it competes against. According to figures published by Inherent itself, Faraday won 73 percent of tasks resembling its training data, and stayed ahead with 60 percent of wins on entirely new, unseen papers. Hughes stresses that beating the competing models mattered less to Inherent than the method behind it: Faraday was trained via reinforcement learning, which rewards good outcomes rather than dictating exact procedures. The goal was to give the system a kind of research instinct, a feel for which experimental approaches are worth pursuing.

The catch: self-measured, self-graded

As impressive as the numbers sound, their reliability remains unclear. Inherent has not named specific test papers nor disclosed its full evaluation methodology, as outlets like Superpower Daily have critically noted. Independent verification of the claims is therefore currently impossible. There is also a methodological wrinkle: Faraday relies on OpenAI’s own Codex model for the actual coding work, meaning the result measures the interplay of multiple systems rather than the performance of a single model. And even the most generous reading of the result only shows that Faraday can reproduce already-published, known outcomes, not that the agent can independently generate new scientific insights. The same caution around self-published AI benchmarks applies broadly, as a recent study on the weaknesses of common AI safety tests also illustrated.

A startup with momentum

Despite the open questions around its evidence, Inherent has already raised capital: a $50 million seed round that is meant to grow the team from roughly a dozen employees today to 20 to 25 by year’s end. For investors, immediate independent confirmation apparently matters less than the underlying bet: smaller, purpose-trained models might be able to match much larger frontier systems on narrowly defined tasks, if training is tailored precisely to that task. If confirmed, that would challenge the pure scaling story pushed by the largest AI labs, in which more parameters and more compute are assumed to reliably produce better results.

Context and outlook

Whether Faraday genuinely marks a turning point for lean, specialized AI research agents, or is simply a well-timed announcement to support its own funding round, cannot be seriously judged yet. What matters is whether Inherent actually makes its Replica benchmark publicly accessible, and whether independent research groups can reproduce the results. Until then, the usual caution applies to self-published comparison numbers from startups mid-fundraise: impressive percentages are no substitute for third-party verification.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top