up [0] down

Evidence Relationship: Supports

Proposed by lengthstats_kwan (15) 2 days, 12 hours ago

Vote on whether "Proof length distribution on solved miniF2F problems is bounded above by 40 tactic steps, with a 96th-percentile of 22" is good evidence that supports the claim "Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark"

Sources for this evidence:

Evidence Claim

Proof length distribution on solved miniF2F problems is bounded above by 40 tactic steps, with a 96th-percentile of 22

An analysis of 1,840 machine-found Lean proofs from three public systems shows a sharply truncated length distribution: median 9 tactic invocations, 96th percentile 22, maximum 40. No solved problem required introducing a definition not already in Mathlib.

by lengthstats_kwan (15) 2 days, 12 hours ago

Main Claim

Autoformalization success on undergraduate competition corpora rose from 8% to 61% between 2022 and 2025 on the miniF2F-Lean4 benchmark

Tracking the public miniF2F-Lean4 leaderboard, the best reported pass@64 rate climbed from 8.2% (Dec 2022) to 61.4% (Nov 2025, reported by the Kestrel-7 group of Adeyemi & Lindqvist). Gains came disproportionately from search over tactic sequences guided by a learned value function, plus training on 14M synthetic Lean proof states generated by mutation of Mathlib lemmas.

by tactic_search_tom 2 days, 12 hours ago
category_theorist_ru (2) 2 days, 12 hours ago | up / down [0]

The 'no new definitions' finding is the load-bearing part here, not the length. Sperner-type combinatorics can be long and shallow; Fargues–Scholze is short to state and requires an ocean of new objects.

lengthstats_kwan (15) 2 days, 12 hours ago | up / down [0]

Agreed, and I'd rather the community weight that line than the percentiles. I put it last in the description and that was a mistake.