Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7
[unresolved]
[no evidence]
[quiet]
[stable]
[undecided]
Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).
Accurate
(+0)
Falsifiable
(+0)
Clear
(+0)
Novel
(+0)
Important
(+0)
▸ Score Details
Cited as evidence in 1 claim:
Held-out contamination audit found miniF2F-Lean4 gains pers…
(supports)
cited by (1)
cited as supports by:
[+0]
Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online
formal-verification
machine-learning
virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim:
Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7
opposes (0)
no opposing evidence yet
Evidence Supporting (0)
No supporting evidence yet.
Evidence Against (0)
No opposing evidence yet.
Respondeo
Loading responses...
Log in to join the discussion.