Evidence receipt / disagreement
Published · transcript-backedOlivia Watkins: disagreement
23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
“Progress has kind of stalled. And basically we realized that this is because the eval is effectively saturated and also highly contaminated.”
Source trail
Everything needed to verify it.
- Speaker
- Olivia Watkins
- Attribution
- Verified speaker
- Claim type
- disagreement
- Recorded
- 23 Feb 2026
- Publisher
- Latent Space
Transcript context
…So you’ve, you’ve seen the evolution of coding benchmarks over time, and I I think it was round about to the mid to late 2024 when you first covered three verified. These have evolved a lot since then. What’s the blog post that you have worked on that you, that we’re releasing today? Like what, what is the sort of con, what’s the main thesis that you’re pushing out? So the main thesis is that SWE-Bench verified has been one of the North Star coding benchmarks that the field has looked at to measure coding progress. But recently we’ve seen that. Progress has kind of stalled. And basically we realized that this is because the eval is effectively saturated and also highly contaminated. So at this point we think that it’s not really measuring coding performance improvements well anymore. And we think that the field should move away from this towards other benchmarks like SWE-Bench Pro.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.