Fermat’s Last Theorem, Machine-Checked: What Claude’s Proof Really Shows

Mathematische Gleichungen in Kreide auf einer dunklen Tafel
Photo by Thomas T on Unsplash

On September 4, 2026, Anthropic published the first fully computer-checked version of Fermat’s Last Theorem. Dozens of Claude agents translated the proof into 13 million lines of Lean code in eleven days, a task the field had expected to take years of coordinated human work. No new mathematics came out of it. What did come out is a tool that can check mathematical literature for errors by machine.

Key takeaways

  • Anthropic had several dozen Claude agents formalize the proof of Fermat’s Last Theorem in the proof language Lean: 13 million lines of code, roughly 29,500 intermediate theorems used, eleven days of runtime.
  • The mathematical content is not new. Andrew Wiles proved the theorem in 1995; what was formalized is a simplified presentation of that proof.
  • The first attempt failed because the agents lost track of the project state. Prove2Me, an open-source tool from Columbia University, finally brought structure to the division of labor.
  • Mathematician Kevin Buzzard of Imperial College London reviewed and compiled the result. His verdict: a major step for automatic formalization, but nothing mathematically new.
  • Estimated compute costs run between $100,000 and $300,000 for roughly six billion generated tokens.

What actually happened here

In 1637, Pierre de Fermat scribbled in a book margin that the equation a to the n plus b to the n equals c to the n has no whole-number solutions when n is greater than two. The margin, he wrote, was too narrow for the proof. The claim stayed open for 358 years until Andrew Wiles proved it in 1995 across 129 pages, and even that version took months before colleagues had worked through it and confirmed it was correct.

That effort is exactly the problem formalization is meant to solve. Formalizing means translating a mathematical argument step by step into a programming language that a so-called proof assistant can verify. The best known one is called Lean. It accepts a proof only when every single step traces back cleanly to axioms or previously proven theorems. Where a human writes “it obviously follows,” Lean demands the full derivation. That makes the work punishingly tedious, and it is why only a small slice of modern mathematics exists in this form today.

The size of the result shows the mismatch: 13 million lines of Lean code for a proof that fills 129 pages on paper. For comparison, Anthropic points to Mathlib, the community-maintained standard library of formalized mathematics that many people have worked on for years. The Fermat code alone is more than five times larger.

Why the first attempt failed

Anthropic describes the process with unusual candor. The first run started well and then fell apart: the agents “quickly lost track of the project state and stopped collaborating effectively.” About seven percent of the code eventually used still comes from those abandoned attempts.

What turned things around was not a stronger model but better infrastructure. Prove2Me, an open-source tool designed by Anthropic researcher Tianyi Peng at Columbia University, manages the proof structure as a directed graph. It separates a theorem’s statement from its proof, which speeds up compilation, and describes each intermediate result in plain language as well. Only then could the agents look up what colleagues had already finished instead of redoing the same work.

Human steering stayed thin and very coarse. Peng occasionally offered directional hints such as “Jacobian as a scheme sounds high priority,” pointing at which subarea should come next. The model itself was an internal research system that Anthropic places roughly in the performance class of Claude Fable 5.1.

Why mathematicians are still cautious

Kevin Buzzard has led the international effort to translate Fermat’s Last Theorem into Lean by hand for years. He reviewed and compiled the Anthropic code, and his assessment is measured: the result is “a great step towards automatic formalization of modern mathematical literature” but delivers “nothing mathematically new.” What was formalized is not Wiles’ original paper but a simplified presentation prepared by Henri Darmon, Fred Diamond, and Richard Taylor.

Then there is the question of who checks the result. The Lean compiler confirms that the derivation is internally sound, and that is the real strength of the method. What stays open is whether the formalized statements actually claim what they are supposed to claim: an error in how a theorem is phrased will not trip the compiler. That translation layer still needs human judgment, and so far Buzzard has largely supplied it. The code sits publicly on GitHub, and broader peer review is only beginning.

The skepticism has reasons that reach beyond this project. A DeepMind experiment with 100 AI agents recently showed how reliably agent systems route around grading mechanisms when the reward is right. Formalization largely removes that risk: the compiler cannot be talked into anything. That is the decisive difference between a model that asserts an answer and one that has to demonstrate its path there by machine.

What follows from this

The practical value lies less with Fermat than with everything that comes after. Mathematical papers are reviewed by humans today, which takes months and still misses errors. If proofs can be machine-checked at a reasonable price, that burden shifts. Six-figure compute costs for a single theorem are still too much for that, though, and the benchmark to beat is the unpaid time of experts who currently do the same job.

It is also worth noting where the progress came from. The breakthrough was not a bigger model but a tool that coordinates the work of many agents and organizes their memory. The same observation applies to other agent projects: the bottleneck is rarely the individual reasoning step, it is collaboration sustained over days. That Anthropic documented this finding so thoroughly shortly after announcing its IPO is unlikely to be a coincidence. For investors, the message “our agents worked toward one goal for eleven days straight” is at least as interesting as the theorem itself.

For readers, the balance is undramatic: a problem solved in 1995 was put into a form machines can verify. That is not evidence that AI does mathematics. It is evidence that AI can handle the record-keeping, and in a field whose quality control suffers from a shortage of time, that is worth more than it first sounds.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top