Evidence receipt / belief
Published · transcript-backedErik Schluntz: belief
28 Nov 2024 Latent Space The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
“From the day-to-day job. But I think one of the most interesting things about SWE-Bench is that all these other benchmarks are usually just isolated puzzles, and you're starting from scratch.”
Source trail
Everything needed to verify it.
- Speaker
- Erik Schluntz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Nov 2024
- Publisher
- Latent Space
Transcript context
…Dynamic programming. Yeah, exactly. Like, you code. From the day-to-day job. But I think one of the most interesting things about SWE-Bench is that all these other benchmarks are usually just isolated puzzles, and you're starting from scratch. Whereas SWE-Bench, you're starting in the context of an entire repository. And so it adds this entirely new dimension to the problem of finding the relevant files. And this is a huge part of real engineering, is it's actually pretty rare that you're starting something totally greenfield. You need to go and figure out where in a codebase you're going to make a change and understand how your work is going to interact with the rest of the systems. And I think SWE-Bench does a really good job of presenting that problem. Why do we still use human eval? It's like 92%, I think. I don't even know if you can actually get to 100% because some of the data is not actually…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.