Evidence receipt / evaluation
Published · transcript-backedMia Glaese: evaluation
23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
“Where, where that’s a significant, it like that means that if you pass a test, actually, like you probably did like a really good job, but just because you didn’t pass a VE test [00:08:00] doesn’t mean that your implementation wasn’t like a good one.”
Source trail
Everything needed to verify it.
- Speaker
- Mia Glaese
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Feb 2026
- Publisher
- Latent Space
Transcript context
…And another set of types of bad tests or tests that are just looking for additional features that were never mentioned. The problem description. Where, where that’s a significant, it like that means that if you pass a test, actually, like you probably did like a really good job, but just because you didn’t pass a VE test [00:08:00] doesn’t mean that your implementation wasn’t like a good one. Right. So it was just like, we only accept like very narrow versions of solutions and like not the whole space Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.