High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

13 Dec 2024 Latent Space Windsurf: The Enterprise AI IDE - with Varun and Anshul of Codeium AI

“What really matters here is like across a broad set of tasks, you're performing like high quality sort of suggestions for people and people love using the product. And I think actually like the way these things work is beyond a certain point, because yes, I actually think it's valuable beyond a certain point.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
13 Dec 2024
Publisher
Latent Space

Transcript context

…od first pass. And I'm not saying it's perfect, but it's only going to keep getting better. And we have deep infrastructure to that actually is validating that we are getting better on this dimension. You mentioned the end-to-end evals that we have for the system, which I think are super cool. But I think you can even decompose each of those steps, right? The ideas of just take retrieval, for example. How can we make eval for retrieval really good? And I think this is just a general thing that's been true about us as a company. It's like most evals and benchmarks, that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have Swybench, that's cool. No actual professional work looks like Swybench. Like, human eval, same thing. Like, these things are just a little kind of broken. So when you're trying to optimize against a metric that's a little bit broken, you end up making kind of suboptimal decisions. So something that we're always very keen on is like, okay, what is the actual metric that we want to test for this part of the system? And so take retrieval, for example. A lot of the benchmarks for these embedding-based systems are like needle in the haystack problems. Like, I want to find this one particular piece of information out of all this potential context. That's not really what actually is necessary for doing software engineering because code is a super distributed knowledge store. You actually want to pull in snippets from a lot of different parts of the code base in order to do the work, right? And so we built systems that, instead of looking at retrieval at one, you're looking at retrieval at like 50. What are the 50 highest things that you can actually retrieve? And are you capturing all of the necessary pieces for that? And what are all the necessary pieces? Well, you can look again back at old, old commits and see what were all the different files that together were edited to make a commit because those are semantically similar things that might not actually show if you actually try to map out a code graph, right? And so we can actually build these kind of golden sets. We can do this evaluation even for sub-problems in the overall task. And so now we have like, you know, an engineering team that can iterate on all of these things and still make sure that the end goal that we're trying to build to is like really, really strong so that we have confidence of what we're pushing out. ing team that can iterate on all of these things and still make sure that the end goal that we're trying to build to is like really, really strong so that we have confidence of what we're pushing out. And by the way, just to talk, we'll say one more thing about the sweep bench thing. Just to showcase these existing, I think benchmarks are not a bad thing. You do want benchmarks. Actually, like I would prefer if there are benchmarks versus let's say everything was just vibes, right? But vibes are also very important, by the way, because they showcase that where the benchmark is not valuable because actually vibes sometimes show you where criminal issues are sort of exist in the benchmark. But like you look at some of the ways in which people have like optimized sweep bench, it's like make sure to run PyTest every time X happens. And it's like, yeah, like sure. You can start like prompting it in like every single possible way. And like, if you remove that, suddenly it doesn't get good at it. It's like, what really matters? What really matters here? What really matters here is like across a broad set of tasks, you're performing like high quality sort of suggestions for people and people love using the product. And I think actually like the way these things work is beyond a certain point, because yes, I actually think it's valuable beyond a certain point. But once it starts hitting the peak of these benchmarks, getting that last 10% actually probably is like counterintuitive to the actual goal of what the benchmark was. Like you probably should find a new hill to climb rather than sort of p-hacking or really optimizing for how you can get higher on the benchmark. Yeah. We did an episode with Anthropic about their recent, like SwyAgent, SwyBench results. And we talked about the human eval versus SwyBench. And like human eval is kind of like a Greenfield benchmark. You know, you need to be good at that. SwyBench is more existing, but it sounds like, I mean, your eval creation is similar to SwyBench as far as like using GitHub commits and kind of like that history. But then it's more like masking at the commit level versus just testing the output of the, of the thing. Cool.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence