Evidence receipt / evaluation
Published · transcript-backedOlivia Watkins: evaluation
23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
“I think the most common problem are like overly narrow tests where there’s some particular implementation detail that the tests we’re looking for, but wasn’t specified in the problem description.”
Source trail
Everything needed to verify it.
- Speaker
- Olivia Watkins
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Feb 2026
- Publisher
- Latent Space
Transcript context
…Yeah. And you found some. Oh, yes. Like in over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are like overly narrow tests where there’s some particular implementation detail that the tests we’re looking for, but wasn’t specified in the problem description. So it wasn’t fair to expect that model to make that particular design choice like one. Pretty blatant example are cases where the task asks you to implement some feature and the tests are looking for you naming that argument or that function with a particular name. But if you may chose another reasonable name, the test would fail. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.