High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Scott Wu: evaluation

4 May 2025 Lenny's Podcast Inside Devin: The world’s first autonomous AI engineer that's set to write 50% of its company’s code by end of year | Scott Wu (CEO and co-founder of Cognition)

“I mean, there's so much detail and idiosyncrasy to the work that we all do obviously day to day. And a lot of it is kind of like teaching the model to mirror the complexity of the real world, I would say, rather than getting it to some higher fundamental level of problem solving, which I think the foundation labs are doing a really great job.”

— Scott Wu

Source trail

Everything needed to verify it.

Speaker
Scott Wu
Attribution
Verified speaker
Claim type
evaluation
Recorded
4 May 2025
Publisher
Lenny's Podcast

Transcript context

…Awesome. I'm going to spend a little time on the tech that enables Devin. Without divulging trade secrets, just what allowed you to make Devin so good? Was there an unlock with a certain model? Some folks have shared three points on a 3.5 was a huge unlock for a lot of their products. Just what's kind of the key to the way you've architected or built Devin that makes it work so well? We obviously, we've been betting on agents for a long time. I think that agents were doable and workable a lot earlier than most folks might've thought. But certainly I think as the community has really rallied around it, I mean you see the impacts of that in the pre-training, you see the impacts of that in a lot of the work that's done with these models. I actually don't think there's been any, from our perspective, I don't think there's been any single step function based model shift or anything that has been kind of like a night and day difference in Devin, but I certainly think that the curve of every point on the chart, I mean there's a new model that comes out every week now has obviously made a big difference in terms of what we've been able to do. And then obviously on top of that, we work with the research teams at all these foundation labs to do a lot of our work on top. And so I think that my hot take here at which I would give is, I think in terms of base intelligence, we're honestly basically already there. And I think a lot of what we see actually and what we spend our time on is less so, obviously, we don't our own models or things like that. It's less so increasing the base IQ of a model, for example, and more about teaching it all of the idiosyncrasies of real-world engineering and thinking about here's how you use Datadog and do this, and here's how you might diagnose this error and here are the different things that you could run into and here's how you handle each of those. And when you're ready, here's how you make GitHub PR. And this is true in engineering. It's true in every other space as well. I mean, there's so much detail and idiosyncrasy to the work that we all do obviously day to day. And a lot of it is kind of like teaching the model to mirror the complexity of the real world, I would say, rather than getting it to some higher fundamental level of problem solving, which I think the foundation labs are doing a really great job. There's something you shared when we were chatting before we started recording around the growth of previous transformative technologies were very hardware oriented and there was a limiting factor to their growth and AI is not that. Can you just share that insight?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence