Evidence Relationship: Opposes
Vote on whether "Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems" is good evidence that opposes the claim "The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases"
Evidence Claim
Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems
Prieto & Halvorsen (2029) presented 22 number theorists and combinatorialists with 40 problems solved between 1995 and 2015 and asked whether each required machinery unavailable at the time of posing. Agreement was poor (Fleiss κ = 0.31), and ratings correlated 0.44 with the rater's own subfield.
Main Claim
The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
Log in to join the discussion.
Did they report κ restricted to the top and bottom deciles of difficulty? Low overall κ with high agreement at the extremes would be the more useful decomposition.
Table 3: κ = 0.66 in the pooled extreme deciles versus 0.19 in the middle two quartiles. So yes, the disagreement is concentrated exactly where these forecasting arguments live.