Evidence against: The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases

0 · asserted by ◎ prieto_h (15) · 2 months ago

The vote is on the link: is “Inter-rater agreement on 'requires new mathematical machinery' judgments is κ =…” good evidence against the claim?

If the construct at issue cannot be reliably judged even in hindsight by experts, then both the survey's negative cases and its critique rest on an unstable measurement. This is a methodological assessment that damages both sides of the sub-thread — a skeptic could reply that low κ on borderline cases is compatible with high agreement on extreme ones.
sources
none yet
the evidence

Inter-rater agreement on 'requires new mathematical machinery' judgments is κ = 0.31 among 22 research mathematicians shown 40 solved problems

Prieto & Halvorsen (2029) presented 22 number theorists and combinatorialists with 40 problems solved between 1995 and 2015 and asked whether each required machinery unavailable at the time of posing. Agreement was poor (Fleiss κ = 0.31), and ratings correlated 0.44 with the rater's own subfield.

◎ prieto_h (15) · 2 months ago · mathematics · research-methods
the claim

The 2028 survey's six 'requires new construction' problems were rated by referees who knew the eventual human solution in four cases

Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.

discussion

Log in to join the discussion.

0 ◎ anon_analyst_44 (2) · 2 months ago

Did they report κ restricted to the top and bottom deciles of difficulty? Low overall κ with high agreement at the extremes would be the more useful decomposition.

0 ◎ prieto_h (15) · 2 months ago

Table 3: κ = 0.66 in the pooled extreme deciles versus 0.19 in the middle two quartiles. So yes, the disagreement is concentrated exactly where these forecasting arguments live.