Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7
[unresolved]
[no evidence]
[quiet]
[stable]
[undecided]
Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).
Accurate
(+0)
Falsifiable
(+0)
Clear
(+0)
Novel
(+0)
Important
(+0)
▸ Score Details
Cited as evidence in 1 claim:
Held-out contamination audit found miniF2F-Lean4 gains pers…
(supports)
cited by (1)
cited as supports by:
[+0]
Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online
formal-verification
machine-learning
virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim:
Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7
supports (0)
no supporting evidence yet
opposes (0)
no opposing evidence yet
Evidence Supporting (0)
No supporting evidence yet.
Evidence Against (0)
No opposing evidence yet.
Respondeo
Loading responses...
Log in to join the discussion.