High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Laura Burkhauser: evaluation

6 May 2026 The Cognitive Revolution "Descript Isn't a Slop Machine": Laura Burkhauser on the AI Tools Creators Love and Hate

“When it comes to the default, that is where we do do, we look at external evals, and then we run some of our own on common customer use cases to find out, like, generally we think that people are going to have the best experience.”

— Laura Burkhauser

Source trail

Everything needed to verify it.

Speaker
Laura Burkhauser
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 May 2026
Publisher
The Cognitive Revolution

Transcript context

…I love the emphasis on play. That's one of my. most common refrains as well. This technology rewards play, and not just the video or visual generative models, but really all of the current frontier AI capabilities really reward play more than any other technology I've ever used. And that really is the right mindset to go into it. And I couldn't agree more with that. I think there's like 6 different follow-up questions that I want to ask based on everything that you just told me. And maybe the first one would be, How do you choose which generative models to put into a product? There are obviously many. They have very different strengths and weaknesses. They have different price points. And there's another question coming out. price and how you're thinking about that and managing that. But these things are super hard to benchmark, right? It's not like in a, when we get to the underlord portion, I think you'll have a much clearer line of sight to like, is this, you know, new model or new prompt or whatever, like doing what we want it to do in a reliable way for a finite set of understood use cases. With the generative stuff, it's tough. Is it just vibes or do you have a better answer for like how you're figuring out what to actually pull the trigger on moving into the product? Yeah, so there's sort of two stage gates. The first is like, should this be available within Descript? And the second is, should we make this the default model? Because most people are not going to change the default model. They're going to accept whatever you put as the default model, right? That actually might be surprising to you. I feel like that's something that if you're deep in AI, you're sort of like, why aren't you using the model picker? Obviously, Nano Banana 2 Pro is going to be like the best thing for photorealistic like face swaps, but then you should be using Cling for this other usage, right? That's how people who are deep in think about things, but the average kind of person doesn't have that level of sophistication, doesn't want that level of sophistication. And so like, how do we make decisions about default models? How do we make decisions about what models to improve or to bring in is a little bit vibes. I'm not gonna lie, because it's not like we like eval every single model out there and say like, these are the five best or whatever. So it's often things where It needs to be available via kind of the, we use FAL as our provider. And so if you're not in FAL, you're not going to be in Descript because we don't want to build our own custom connector for your thing unless it's like the best thing ever. But that means we need to like sign a new data license agreement and all this stuff that's like, what a headache. We've already done it with FAL. We're just going to do it there. And so that's why like SeaDance is now in Descript is it is finally in FAL. So we're like, great, you can come on in. And then within stuff that's in file, we try to pick the stuff that feels like generally the best or in the game. Because what you see are these standard kind of industry benchmarks of these different things. And you'll sort of like see that you have the same labs on the leaderboard kind of month after month. And so we try to make sure that we have some representation from each of those labs because you're always like one week away from that lab coming back to the top and having the best thing available. When it comes to the default, that is where we do do, we look at external evals, and then we run some of our own on common customer use cases to find out, like, generally we think that people are going to have the best experience. Now we have, like, for image generation, Nano Banana Pro, I think is our new default. And what we then do is we'll AB test it against the existing default and make sure that we're seeing kind of good things from the AB test and that the AB test matches kind of like what our internal evals tell us. And if it does, then it's like a definite ship. This is our new default. When you do an internal eval, is it a panel of trusted people that are like scoring outputs?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence