Evidence Relationship: Supports
Vote on whether "Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online" is good evidence that supports the claim "Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark"
Evidence Claim
Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online
Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.
Main Claim
Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark
Tracking the public miniF2F-Lean4 leaderboard, the best reported pass@64 rate climbed from 8.2% (Dec 2022) to 61.4% (Nov 2025, reported by the Kestrel-7 group of Adeyemi & Lindqvist). Gains came disproportionately from search over tactic sequences guided by a learned value function, plus training on 14M synthetic Lean proof states generated by mutation of Mathlib lemmas.
Log in to join the discussion.
This is the sub-evidence that makes the parent worth taking seriously. Contamination is the default explanation for every benchmark jump and someone actually tested it.