AI

An AI wrote 13 million lines of maths in eleven days. A compiler checked every one of them.

Anthropic says Claude produced the first end-to-end machine-checked proof of Fermat’s Last Theorem in Lean 4 — roughly 30,300 theorems, about six billion output tokens, largely unsupervised. Mathematicians had budgeted years. The speed is not the point.

N Noah · The Sharp Brief · September 6, 2026 · 5 min read

Anthropic said this week that Claude produced the first complete, machine-checked proof of Fermat’s Last Theorem in Lean 4 — the formal proof language mathematicians use when they want a computer, not a referee, to sign off on the argument. The run took about eleven days. It generated roughly 13 million lines of Lean, proved some 30,300 theorems — around 29,500 of which the final proof actually leans on — and burned through approximately six billion output tokens.

The work was assembled on Prove2Me, an open platform for collaborative formalisation built by Tianyi Peng and collaborators at Columbia. Human mathematical input was limited to occasional high-level direction. Anthropic describes the model behind it as a general-purpose internal research system roughly comparable to Claude Fable 5.1 — not a bespoke theorem-proving machine.

Formalising Andrew Wiles’ 1994 proof has been an open project for years. The estimates for doing it by hand ran into the multiple-year range. That gap — years to under two weeks — is the headline. It is not the interesting part.

The compiler is the story

Every one of those 13 million lines had to pass the Lean kernel. A formal proof assistant does not grade on effort or plausibility. A line either type-checks or it does not. Which means the failure mode that has defined two years of enterprise AI anxiety — the model produces something confident, fluent and wrong — is structurally impossible here. Hallucinated mathematics simply does not compile.

That is the actual result. Not “the model is smarter,” but “when you point a model at a domain that has a mechanical verifier, you can let it run for eleven days unsupervised and trust the output.” Six billion tokens of unreviewed generation would be an unusable liability in almost any other setting. Behind a checker, it is just throughput.

Our take: the bottleneck in applied AI has never really been generation. It is verification — the human hours spent confirming that what came out is true. Fermat is a demonstration of what happens when you remove that cost entirely. The lesson for anyone deploying agents is not “buy more model.” It is: find the parts of your workflow that have a ground-truth oracle, and put the agents there first.

Where this transfers, and where it does not

Formal mathematics is the extreme case, but it is not the only domain with a mechanical checker. Code that has to pass a test suite. SQL that has to return against a real schema. Financial reconciliations that have to tie to a control total. Configuration that has to survive a linter and a staging deploy. In all of those, the model can be wrong repeatedly and cheaply, because something downstream catches it before a human does.

Contrast that with the places most companies have actually deployed AI: drafting, summarising, research memos, customer replies. Nothing checks those. The verification cost lands on a person, and it does not fall as the model improves — it arguably rises, because more plausible output is harder to audit. That is why so many pilots stall at “impressive demo, unclear savings.”

Two caveats worth holding. First, this is Anthropic reporting on its own model, and the wider mathematical community has not finished picking the artefact apart. The proof object is checkable by anyone, which is a far stronger position than a benchmark score — but “checkable” and “checked by others” are different states. Second, the run consumed roughly six billion output tokens. Whatever this cost, it was not cheap, and the economics of pointing that much compute at a problem only work when the answer is worth it.

What to watch

The uncomfortable read for most AI buyers: the biggest unlock of the last month was not a new model. It was a reminder that models get dramatically more useful the instant you can stop reading their homework.

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.