Evidence receipt / belief
Published · transcript-backedErik Schluntz: belief
28 Nov 2024 Latent Space The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
“I think there's still need for sort of many different varied evals. Like sometimes you do really care about just sort of greenfield code generation.”
Source trail
Everything needed to verify it.
- Speaker
- Erik Schluntz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Nov 2024
- Publisher
- Latent Space
Transcript context
…Do you see benchmarks like that, they should just get sunsetted? Because when you look at the model releases, it's like, oh, it's like 92% instead of like 89%, 90% on human eval versus, you know, SWE-Bench verified is you have 49%, right? Which is like, before 45% was state of the art, but maybe like six months ago it was like 30%, something like that. So is that a benchmark that you think is going to replace human eval, or do you think they're just going to run in parallel? I think there's still need for sort of many different varied evals. Like sometimes you do really care about just sort of greenfield code generation. And so I don't think that everything needs to go to sort of an agentic setup. It would be very expensive to implement.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.