High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Shawn Wang: belief

23 Feb 2026 Latent Space ⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data

“I think the let’s, let’s sort of reset on like, what was the original work that you guys did for Verified, which I think was pretty substantial.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
belief
Recorded
23 Feb 2026
Publisher
Latent Space

Transcript context

…Like SWE-Bench Pro. Yeah. Amazing. Yeah, I, one of the jokes I always have is like there’s a group chat with all the labs. And everyone just takes turns, the increment, like 0.1 on trucks and then it’s like, okay, well you have the best coding model, I guess. ‘cause you’re 0.1% higher, but it’s not super convincing at this point. No. Yeah. So cool. I think the let’s, let’s sort of reset on like, what was the original work that you guys did for Verified, which I think was pretty substantial. Like, it was like a very significant investment from OpenAI, which like people still don’t appreciate. And then. What were the satisfactions that we, that we found over time? Right? So like what, what was Sweet Bench Verify, or should, should that people should know about? Suite bench verified was kind of a cleanup of original bench academic benchmark from a lab at Princeton called Suite Bench. And the agent is basically given a code base and a task that was sourced from a real world repository and GitHub issue, and was asked to solve a task and is graded on whether some tests pass. And at the time this was quickly became a popular benchmark because at the time the field didn’t really have good real world coding benchmarks. Then when open, I took a look at the benchmark as part of one of the evals we wanted to track in our preparedness framework. Folks started realizing that some of the cases where agents were failing were due to bad problem setups rather than just to models being dumb. So folks at OpenAI did a pretty extensive human data campaign hiring like almost a hundred real world software engineers to go through the problems and figure out like, are the tasks well specified? Are the tests actually fair? And kind of created a curated set of like 500 tasks that we thought were much better.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence