High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / disagreement

Published · transcript-backed

Olivia Watkins: disagreement

23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data

“Progress has kind of stalled. And basically we realized that this is because the eval is effectively saturated and also highly contaminated.”

— Olivia Watkins

Source trail

Everything needed to verify it.

Speaker
Olivia Watkins
Attribution
Verified speaker
Claim type
disagreement
Recorded
23 Feb 2026
Publisher
Latent Space

Transcript context

…So you’ve, you’ve seen the evolution of coding benchmarks over time, and I I think it was round about to the mid to late 2024 when you first covered three verified. These have evolved a lot since then. What’s the blog post that you have worked on that you, that we’re releasing today? Like what, what is the sort of con, what’s the main thesis that you’re pushing out? So the main thesis is that SWE-Bench verified has been one of the North Star coding benchmarks that the field has looked at to measure coding progress. But recently we’ve seen that. Progress has kind of stalled. And basically we realized that this is because the eval is effectively saturated and also highly contaminated. So at this point we think that it’s not really measuring coding performance improvements well anymore. And we think that the field should move away from this towards other benchmarks like SWE-Bench Pro.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence