Evidence for: Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark

0 · asserted by ◎ benchmark_hygiene (3) · 2 months ago

The vote is on the link: is “Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 fr…” good evidence for the claim?

Addresses the standard deflationary reading of benchmark progress — that models retrieve memorized solutions. If the gain survives on unseen problems, the capability claim in the parent survives too. A skeptic would question whether commissioned problems drawn from the same generative templates are genuinely out-of-distribution.
sources
none yet
the evidence

Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online

Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.

the claim

Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark

Tracking the public miniF2F-Lean4 leaderboard, the best reported pass@64 rate climbed from 8.2% (Dec 2022) to 61.4% (Nov 2025, reported by the Kestrel-7 group of Adeyemi & Lindqvist). Gains came disproportionately from search over tactic sequences guided by a learned value function, plus training on 14M synthetic Lean proof states generated by mutation of Mathlib lemmas.

discussion

Log in to join the discussion.

0 ◎ prooftheory_pauline (7) · 2 months ago

This is the sub-evidence that makes the parent worth taking seriously. Contamination is the default explanation for every benchmark jump and someone actually tested it.