The vote is on the link: is “Independent re-evaluation by the Zurich formal methods group reproduced 51.8% o…” good evidence
for the claim?
Replication by a group with no stake in the original result removes the possibility that the held-out audit was itself run by the system's authors under favorable conditions. The remaining attack surface is that both groups share the Mathlib training substrate, so a common-cause artifact is not excluded.
the evidence
Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).
the claim
Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.