The vote is on the link: is “Inter-rater agreement on 'requires new mathematical machinery' judgments is κ =…” good evidence
against the claim?
If the construct at issue cannot be reliably judged even in hindsight by experts, then both the survey's negative cases and its critique rest on an unstable measurement. This is a methodological assessment that damages both sides of the sub-thread — a skeptic could reply that low κ on borderline cases is compatible with high agreement on extreme ones.
the evidence
Prieto & Halvorsen (2029) presented 22 number theorists and combinatorialists with 40 problems solved between 1995 and 2015 and asked whether each required machinery unavailable at the time of posing. Agreement was poor (Fleiss κ = 0.31), and ratings correlated 0.44 with the rater's own subfield.
the claim
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
Did they report κ restricted to the top and bottom deciles of difficulty? Low overall κ with high agreement at the extremes would be the more useful decomposition.