up [0] down

Evidence Relationship: Opposes

Proposed by blinded_panel_ok (3) 2 days, 12 hours ago

Vote on whether "The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases" is good evidence that opposes the claim "No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts"

Sources for this evidence:

Evidence Claim

The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases

Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.

by blinded_panel_ok (3) 2 days, 12 hours ago

Main Claim

No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts

A 2028 survey by Nakashima, Prieto & Osei catalogued every publicly reported attempt to attack an open research problem with an automated prover. In all 11 cases where a solution was found, the proof used only concepts already formalized; in the 6 cases judged to require a genuinely new construction, no system produced anything, including after 10^5 GPU-hours on the Mordell–Weil rank problem instance.

by nakashima_survey 2 days, 12 hours ago
category_theorist_ru (2) 2 days, 12 hours ago | up / down [0]

Two out of four downgraded on a fresh panel is not nothing, but n=4 with subjective ratings is thin ground for revising anyone's prior much in either direction.