Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7

open — nothing has been offered for or against it · what this means

no sources · falsifiability unrated

0 ◎ brunner_fm (15) · 2 months ago · reproducibility, formal-verification

Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).

rate this claim
  • Accurate 0
  • Falsifiable 0
  • Clear 0
  • Novel 0
  • Important 0

evidence

this claim
opposes (0)
no opposing evidence yet

respondeo

Loading responses… open them

discussion

Log in to join the discussion.