A wave of recent results suggests the latest large language models are cracking mathematical problems that stumped human researchers for decades, prompting claims that AI has crossed into "superhuman" territory even as leading mathematicians urge caution.
June 2026 · Frontier AI in Mathematics
AI Is Cracking Math Problems That Stumped Humans for Decades
Frontier language models are now credited with settling open problems — some verified in Lean — as benchmark scores climb. Leading mathematicians still urge caution: impressive, but unproven at scale.
148 min
To settle a convex-optimization lower bound open since 1996 — proof checked in Lean.
<60 min
To close the Cycle Double Cover Conjecture (a 1970s problem) using dozens of parallel sub-agents.
~38%
GPT-5.4 on FrontierMath — well ahead of rivals below 23%.
FrontierMath — the hardest problem set
GPT-5.4 pulls clear of the field
What changed — three capabilities converged
+
+
Formal verification (Lean)
Together they narrowed a gap once assumed permanent — and compressed expert forecasts: a publishable AI theorem, once pegged near mid-century , and a Putnam-level win, once expected around 2033 .
The optimistic read
A rapid succession of solved open problems suggests AI is becoming a genuine engine of mathematical discovery — with Lean offering a mechanical guarantee that proofs are actually correct.
The cautious read
Terence Tao likens current systems to a talented but unreliable grad student whose every step needs checking. Much rests on provider disclosures and one-off demos, not peer review — and June 2026 testing still found tasks where humans hold the edge.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…