up [0] down

Evidence Relationship: Supports

Proposed by brunner_fm (15) 2 days, 12 hours ago

Vote on whether "Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7" is good evidence that supports the claim "Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online"

Sources for this evidence:

Evidence Claim

Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7

Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).

by brunner_fm (15) 2 days, 12 hours ago

Main Claim

Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online

Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.

by benchmark_hygiene 2 days, 12 hours ago