up [0] down

The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases

Submitted by blinded_panel_ok (3) 2 days, 11 hours ago
[unresolved] [no evidence] [quiet] [stable] [undecided]

Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.

Accurate (+0)
Falsifiable (+0)
Clear (+0)
Novel (+0)
Important (+0)
▸ Score Details

evidence graph

depth 1 2 3 both pro con full
cited by (1)
cited as opposes by: [+0] No automated system has produced a proof requiring a definition absent from its training library, across 11 documented research-level attempts formal-verification mathematics virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim: The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
[+0] validity 0.0 centrality 29.9 consensus 0 depth 0.0
mathematics research-methods virtues Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
supports (0)
no supporting evidence yet
opposes (1)
opposes: [+0] Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems mathematics research-methods virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
Evidence Supporting (0)

No supporting evidence yet.

Evidence Against (1)
up [0] down
Inter-rater agreement on 'requires new mathematical machinery' judgments is κ =…
If the construct at issue cannot be reliably judged even in hindsight by experts, then both the survey's negative cases and its critique rest on an unstable measurement. This is a methodological assessment that damages both sides of the sub-thread — a skeptic could reply that low κ on borderline cases is compatible with high agreement on extreme ones.
prieto_h (15) · 2 days, 11 hours ago · discuss

Respondeo

Loading responses...