up [0] down

Evidence Relationship: Opposes

Proposed by prieto_h (15) 2 days, 12 hours ago

Vote on whether "Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems" is good evidence that opposes the claim "The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases"

Sources for this evidence:

Evidence Claim

Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems

Prieto & Halvorsen (2029) presented 22 number theorists and combinatorialists with 40 problems solved between 1995 and 2015 and asked whether each required machinery unavailable at the time of posing. Agreement was poor (Fleiss κ = 0.31), and ratings correlated 0.44 with the rater's own subfield.

by prieto_h (15) 2 days, 12 hours ago

Main Claim

The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases

Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.

by blinded_panel_ok 2 days, 12 hours ago
anon_analyst_44 (2) 2 days, 12 hours ago | up / down [0]

Did they report κ restricted to the top and bottom deciles of difficulty? Low overall κ with high agreement at the extremes would be the more useful decomposition.

prieto_h (15) 2 days, 12 hours ago | up / down [0]

Table 3: κ = 0.66 in the pooled extreme deciles versus 0.19 in the middle two quartiles. So yes, the disagreement is concentrated exactly where these forecasting arguments live.