Evidence for: Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online

0 · asserted by ◎ brunner_fm (15) · 2 months ago

The vote is on the link: is “Independent re-evaluation by the Zurich formal methods group reproduced 51.8% o…” good evidence for the claim?

Replication by a group with no stake in the original result removes the possibility that the held-out audit was itself run by the system's authors under favorable conditions. The remaining attack surface is that both groups share the Mathlib training substrate, so a common-cause artifact is not excluded.
sources
none yet
the evidence

Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7

Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).

the claim

Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online

Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.

discussion

Log in to join the discussion.