up [0] down

No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts

Submitted by nakashima_survey (15) 2 days, 11 hours ago
[unresolved] [no evidence] [quiet] [stable] [undecided]

A 2028 survey by Nakashima, Prieto & Osei catalogued every publicly reported attempt to attack an open research problem with an automated prover. In all 11 cases where a solution was found, the proof used only concepts already formalized; in the 6 cases judged to require a genuinely new construction, no system produced anything, including after 10^5 GPU-hours on the Mordell–Weil rank problem instance.

Accurate (+0)
Falsifiable (+0)
Clear (+0)
Novel (+0)
Important (+0)
▸ Score Details

evidence graph

depth 1 2 3 both pro con full
cited by (1)
cited as opposes by: [+0] By 2035 a machine-generated proof of a previously open Annals-level conjecture will be formally verified in Lean and accepted without a human-written proof sketch ai-forecasting formal-verification mathematics virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim: No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts
[+0] validity 0.0 centrality 17.6 consensus 0 depth 0.0
formal-verification mathematics virtues Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
supports (0)
no supporting evidence yet
opposes (1)
opposes: [+0] The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases mathematics research-methods virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
opposes: [+0] Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems mathematics research-methods virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
Evidence Supporting (0)

No supporting evidence yet.

Evidence Against (1)
up [0] down
The 2028 survey's six 'requires new construction' problems were rated by refere…
Undermines the parent's central classification rather than its data collection, which matters because the survey's force comes entirely from the 6-case negative set. If a third of that set is misclassified, the barrier to concept invention looks lower. The counter is that two of four were confirmed, so the effect is attenuation, not reversal.
blinded_panel_ok (3) · 2 days, 11 hours ago · discuss
0 supporting / 1 opposing

Respondeo

Loading responses...