The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
[unresolved]
[no evidence]
[quiet]
[stable]
[undecided]
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
Accurate
(+0)
Falsifiable
(+0)
Clear
(+0)
Novel
(+0)
Important
(+0)
▸ Score Details
Cited as evidence in 1 claim:
No automated system has produced a proof requiring a defini…
(opposes)
cited by (1)
cited as opposes by:
[+0]
No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts
formal-verification
mathematics
virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim:
The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
opposes (1)
opposes:
[+0]
Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems
mathematics
research-methods
virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
Evidence Supporting (0)
No supporting evidence yet.
Evidence Against (1)
Inter-rater agreement on 'requires new mathematical machinery' judgments is κ =…
If the construct at issue cannot be reliably judged even in hindsight by experts, then both the survey's negative cases and its critique rest on an unstable measurement. This is a methodological assessment that damages both sides of the sub-thread — a skeptic could reply that low κ on borderline cases is compatible with high agreement on extreme ones.
Respondeo
Loading responses...
Log in to join the discussion.