Evidence Relationship: Opposes
Vote on whether "No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts" is good evidence that opposes the claim "By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch"
Evidence Claim
No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts
A 2028 survey by Nakashima, Prieto & Osei catalogued every publicly reported attempt to attack an open research problem with an automated prover. In all 11 cases where a solution was found, the proof used only concepts already formalized; in the 6 cases judged to require a genuinely new construction, no system produced anything, including after 10^5 GPU-hours on the Mordell–Weil rank problem instance.
Main Claim
By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch
The claim requires four conjuncts to hold jointly before 31 Dec 2035: (a) the target was an open problem of a difficulty class typically published in Annals of Mathematics or comparable venues; (b) the proof's mathematical content originates from an automated system, not from human-supplied intermediate lemmas or a strategy outline; (c) the artifact typechecks in Lean (or a successor with a comparable kernel); (d) the community treats the result as established. Human-written formalization scaffolding of a human-supplied argument does not count.
Log in to join the discussion.
This is the crux of the whole network for me. If concept invention is a distinct capability rather than a continuation of the search curve, the 2035 date is badly optimistic.
I'd push back slightly — 'requires a new definition' is judged post hoc by humans who found the proof that way. Erdős-style combinatorics has repeatedly turned out to need less machinery than expected.
That's a fair limitation and we flag it in §4: the 6 'requires new construction' judgments came from three referees with only moderate agreement (Fleiss κ = 0.58).