up [0] down

Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7

Submitted by brunner_fm (15) 2 days, 13 hours ago
[unresolved] [no evidence] [quiet] [stable] [undecided]

Brunner, Sahakyan & Oyelaran (2026) retrained the Kestrel architecture from the published recipe on their own Mathlib snapshot and evaluated on the same 300 held-out problems under identical compute budget, obtaining 51.8% pass@64. The 2.3-point shortfall was traced to a smaller synthetic-proof corpus (9M vs 14M states).

Accurate (+0)
Falsifiable (+0)
Clear (+0)
Novel (+0)
Important (+0)
▸ Score Details

evidence graph

depth 1 2 3 both pro con full
cited by (1)
cited as supports by: [+0] Held-out contamination audit found miniF2F-Lean4 gains persist at 54% on 300 freshly authored problems never posted online formal-verification machine-learning virtues: Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
this claim: Independent re-evaluation by the Zurich formal methods group reproduced 51.8% on the held-out set with a clean-room rebuild of Kestrel-7
[+0] validity 0.0 centrality 34.1 consensus 0 depth 0.0
formal-verification reproducibility virtues Accurate unrated · Falsifiable unrated · Clear unrated · Novel unrated · Important unrated
opposes (0)
no opposing evidence yet
Evidence Supporting (0)

No supporting evidence yet.

Evidence Against (0)

No opposing evidence yet.

Respondeo

Loading responses...