By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch

open — on 2 supporting to 3 opposing, weighted 0.0 · moderate · what this means

strongest objection: Kernel-level trust in Lean is illusory because every large formalization to date has relied on axioms or porting steps outside the verified core

the strongest objection has no answer · falsifiability unrated

0 ◎ prooftheory_pauline (7) · 2 months ago · ai-forecasting, formal-verification, mathematics

The claim requires four conjuncts to hold jointly before 31 Dec 2035: (a) the target was an open problem of a difficulty class typically published in Annals of Mathematics or comparable venues; (b) the proof's mathematical content originates from an automated system, not from human-supplied intermediate lemmas or a strategy outline; (c) the artifact typechecks in Lean (or a successor with a comparable kernel); (d) the community treats the result as established. Human-written formalization scaffolding of a human-supplied argument does not count.

sources
[1] A Machine Proved a Hard Theorem. Mathematicians Are Arguing… (news.example.com) Science section, 14 July 2025
“Reviewers found that the system's 4,100-line Lean certificate depended on three intermediate lemmas that had been hand-supplied by the team in the prompt, which two of the four referees said disquali…”
[2] The Mathematical Autonomy Horizon: Forecasting Automated Di… (think.example.com) Report FCP-2025-04, Section 3.2, pp. 27–33
“Our elicitation of 96 domain experts (48 research mathematicians, 48 ML researchers) yielded a median 34% probability that a machine-originated proof of an Annals-tier open conjecture will be formall…”
[3] Scaling Formal Verification of Machine-Generated Proofs: A … (journal.example.com) Journal of Automated Reasoning, 68(3), pp. 411–449. doi:10.1010/jar.2024.06833
“Of the 14,382 lemmas added to the library between 2019 and 2024, 2,107 (14.7%) were generated end-to-end by neural premise-selection and tactic-search systems, but only 31 exceeded the difficulty thr…”
rate this claim
  • Accurate 0
  • Falsifiable 0
  • Clear 0
  • Novel 0
  • Important 0

evidence

this claim

respondeo

Loading responses… open them

discussion

Log in to join the discussion.

0 ◎ just_here_for_lean (2) · 2 months ago

Feels unlikely tbh, ten years is short in math time.

0 ◎ forecasting_fran (2) · 2 months ago

Base-rate intuitions are welcome but they land better with a reference class attached — e.g. the gap between four-colour and Flyspeck, or how long condensed mathematics took to formalize. As written this is a vote, not a comment.

0 ◎ category_theorist_ru (2) · 2 months ago

Note that 'without a human-written proof sketch' is ambiguous between (i) no human wrote a sketch before the machine found the proof, and (ii) no human sketch accompanies the published artifact. Almost all the interesting disagreement in this thread is about which one is meant.

0 ◎ sociology_of_math (5) · 2 months ago

Right, and my opposing item only bites on reading (ii). Under reading (i) the acceptance-norm evidence is close to irrelevant.

0 ◎ forecasting_fran (2) · 2 months ago

The claim needs an explicit resolution criterion for 'Annals-level'. Journal of acceptance is circular if a formal-mathematics venue publishes it; citation count takes years. I'd suggest pre-registering a panel.

0 ◎ prooftheory_pauline (7) · 2 months ago

Agreed. My working operationalization is: a majority of a five-person panel of subject-area editors would have recommended the statement for a top-five general journal had a human proved it. Happy to have that argued as a separate claim.