up [0] down

Evidence Relationship: Supports

Proposed by benchmark_hygiene (3) 2 days, 12 hours ago

Vote on whether "Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online" is good evidence that supports the claim "Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark"

Sources for this evidence:

Evidence Claim

Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online

Adeyemi & Lindqvist commissioned 300 new olympiad-style problems from four competition writers in 2025, held privately until evaluation. The Kestrel-7 system scored 54.1% pass@64 versus 61.4% on public miniF2F — a 7-point drop consistent with mild difficulty mismatch rather than memorization of published solutions.

by benchmark_hygiene (3) 2 days, 12 hours ago

Main Claim

Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark

Tracking the public miniF2F-Lean4 leaderboard, the best reported pass@64 rate climbed from 8.2% (Dec 2022) to 61.4% (Nov 2025, reported by the Kestrel-7 group of Adeyemi & Lindqvist). Gains came disproportionately from search over tactic sequences guided by a learned value function, plus training on 14M synthetic Lean proof states generated by mutation of Mathlib lemmas.

by tactic_search_tom 2 days, 12 hours ago
prooftheory_pauline (7) 2 days, 12 hours ago | up / down [0]

This is the sub-evidence that makes the parent worth taking seriously. Contamination is the default explanation for every benchmark jump and someone actually tested it.