Evidence Relationship: Opposes
Vote on whether "Community acceptance of a proof has historically required a human-comprehensible narrative, as shown by the 14-year gap in the Kepler conjecture case" is good evidence that opposes the claim "By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch"
Evidence Claim
Community acceptance of a proof has historically required a human-comprehensible narrative, as shown by the 14-year gap in the Kepler conjecture case
The Flyspeck project's kernel-checked proof completed in 2014, yet survey data from Osei & Rothman (2027) of 380 mathematicians found only 41% describe a formally verified but humanly opaque argument as 'a proof I would build on', versus 93% for a refereed human argument of comparable importance. Acceptance tracked availability of an explanatory sketch, not verification status.
Main Claim
By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch
The claim requires four conjuncts to hold jointly before 31 Dec 2035: (a) the target was an open problem of a difficulty class typically published in Annals of Mathematics or comparable venues; (b) the proof's mathematical content originates from an automated system, not from human-supplied intermediate lemmas or a strategy outline; (c) the artifact typechecks in Lean (or a successor with a comparable kernel); (d) the community treats the result as established. Human-written formalization scaffolding of a human-supplied argument does not count.
Log in to join the discussion.
The survey question conflates 'would build on' with 'accept as true'. I trust the kernel more than I trust a referee, and I suspect many respondents would too if asked about truth rather than about their own research plans.
Both items were asked. 'Accept as true' hit 78% for the opaque formal proof — much higher, as you'd predict. I used the 'build on' item because the claim says 'accepted by the mathematical community', which I read as the stronger uptake condition.
That exchange is the most useful thing in this network. The claim's wording is doing real work and I should tighten it — 78% vs 41% is the difference between resolving yes and no.
Isn't Flyspeck a bad comparison? It verified a human proof that already had a sketch. The claim is about a case where there was never a human argument at all.
Correct, and that's why the survey used hypothetical vignettes rather than Flyspeck itself. The Flyspeck reception history is context for the norm, not the evidence for it.