Evidence receipt / prediction
Published · transcript-backedOlivia Watkins: prediction
23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
“I think, I mean like 90% of the problems are things that were estimated to take like an expert software engineer like less than an hour.”
Source trail
Everything needed to verify it.
- Speaker
- Olivia Watkins
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 23 Feb 2026
- Publisher
- Latent Space
Transcript context
…Yeah. Cool. We’re gonna stop reporting CBE verified. Right? And then SWE-Bench Pro Will will be sort of the next one, which is an effort from scale. What’s your sort of comparison analysis? What’s, what attracts you to SWE-Bench Pro? The first one I think is just that it’s harder for SWE-Bench verified. I think, I mean like 90% of the problems are things that were estimated to take like an expert software engineer like less than an hour. They’re like very well specified, very self-contained, and the SWE-Bench Pro problems are just bigger and harder. There’s much more head rewind that eval because it’s not saturated. Yeah. Like categories of like one to four hours and four plus.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.