Evidence Relationship: Opposes
Vote on whether "The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases" is good evidence that opposes the claim "No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts"
Evidence Claim
The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
Main Claim
No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts
A 2028 survey by Nakashima, Prieto & Osei catalogued every publicly reported attempt to attack an open research problem with an automated prover. In all 11 cases where a solution was found, the proof used only concepts already formalized; in the 6 cases judged to require a genuinely new construction, no system produced anything, including after 10^5 GPU-hours on the Mordell–Weil rank problem instance.
Log in to join the discussion.
Two out of four downgraded on a fresh panel is not nothing, but n=4 with subjective ratings is thin ground for revising anyone's prior much in either direction.