The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
[unresolved]
[no evidence]
[quiet]
[stable]
[undecided]
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
Accurate
(+0)
Falsifiable
(+0)
Clear
(+0)
Novel
(+0)
Important
(+0)
▸ Score Details
Cited as evidence in 1 claim:
No automated system has produced a proof requiring a defini…
(opposes)
cited by (1)
cited as opposes by:
[+0]
No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts
formal-verification
mathematics
virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim:
The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
supports (0)
no supporting evidence yet
Evidence Supporting (0)
No supporting evidence yet.
Evidence Against (1)
Inter-rater agreement on 'requires new mathematical machinery' judgments is κ =…
If the construct at issue cannot be reliably judged even in hindsight by experts, then both the survey's negative cases and its critique rest on an unstable measurement. This is a methodological assessment that damages both sides of the sub-thread — a skeptic could reply that low κ on borderline cases is compatible with high agreement on extreme ones.
Respondeo
Loading responses...
Log in to join the discussion.