The vote is on the link: is “The 2028 survey's six 'requires new construction' problems were rated by refere…” good evidence
against the claim?
Undermines the parent's central classification rather than its data collection, which matters because the survey's force comes entirely from the 6-case negative set. If a third of that set is misclassified, the barrier to concept invention looks lower. The counter is that two of four were confirmed, so the effect is attenuation, not reversal.
the evidence
Follow-up correspondence established that for 4 of the 6 negative cases, at least two of the three referees had read a human proof before rating. Blinded re-rating of those four by a fresh panel downgraded two to 'plausibly solvable with existing machinery'.
the claim
A 2028 survey by Nakashima, Prieto & Osei catalogued every publicly reported attempt to attack an open research problem with an automated prover. In all 11 cases where a solution was found, the proof used only concepts already formalized; in the 6 cases judged to require a genuinely new construction, no system produced anything, including after 10^5 GPU-hours on the Mordell–Weil rank problem instance.
Two out of four downgraded on a fresh panel is not nothing, but n=4 with subjective ratings is thin ground for revising anyone's prior much in either direction.