High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Erik Schluntz: belief

28 Nov 2024 Latent Space The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic

“I think there's still need for sort of many different varied evals. Like sometimes you do really care about just sort of greenfield code generation.”

— Erik Schluntz

Source trail

Everything needed to verify it.

Speaker
Erik Schluntz
Attribution
Verified speaker
Claim type
belief
Recorded
28 Nov 2024
Publisher
Latent Space

Transcript context

…Do you see benchmarks like that, they should just get sunsetted? Because when you look at the model releases, it's like, oh, it's like 92% instead of like 89%, 90% on human eval versus, you know, SWE-Bench verified is you have 49%, right? Which is like, before 45% was state of the art, but maybe like six months ago it was like 30%, something like that. So is that a benchmark that you think is going to replace human eval, or do you think they're just going to run in parallel? I think there's still need for sort of many different varied evals. Like sometimes you do really care about just sort of greenfield code generation. And so I don't think that everything needs to go to sort of an agentic setup. It would be very expensive to implement.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence