Evidence receipt / belief
Published · transcript-backedHowie Liu: belief
31 Aug 2025 Lenny's Podcast How we restructured Airtable’s entire org for AI | Howie Liu (co-founder and CEO)
“You can A/B test, especially if you have the scale of a really large product like Anthropic or OpenAI, you can just test everything and see like, oh, this model actually performs better than this one, this prompt performs better than this one, but I think early on you don't have that luxury and you're in a much more open-ended discovery process.”
Source trail
Everything needed to verify it.
- Speaker
- Howie Liu
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Aug 2025
- Publisher
- Lenny's Podcast
Transcript context
…Yeah, and I listened to your episodes with [inaudible 01:03:54] and Mike who talked about this. I think it's interesting that both heads of OpenAI and Anthropic have converged on this point. I mean, look, I think I would add a slightly different or additive take though, which is I think for a completely novel product experience or form factor, you should actually not start with evals and you should start with vibes, right? Meaning you need to go and just test in a much more open-ended way, like, does this even work in kind of a broad sense? So as an example, for our custom code generation capability, instead of defining evals that get repeatably tested as you vary the prompt or the model or the agentic workflow used to generate these outputs, and you have to define what does good look like by definition for the eval, I would first start with a much more open-ended and ad hoc style of just throw stuff against the wall, try different prompts and see how well it does. To me, evals are more useful, A, once you've converged on the basic scaffold of the form factor and you kind of know what are the use cases you want it to work well for and what you want to test against it. Whereas in the early days, especially if your product market fit finding either for an entirely new company or for a pretty dramatically new or bold new capability that doesn't really have ... it's not an incremental improvement on something that exists in Airtable today, I think you have to just be a little bit more creative initially and throwing stuff at it, seeing what works to understand, okay, let's use an example, we're implementing this new capability that can use basically a long-running AI crawler agent that goes and researches the web for a specific type of object or entity, right? So it's similar to deep research, but what it actually does is instead of outputting a report, it's actually going and compiling a list of things. The things could be companies or people or anything else, right? Find me every Marvel movie ever made, find me every DC Comics spin-off series, literally anything. You have to go in and first just try out a bunch of random ... use your own brain to think of what's the range of use cases I can test this against, right? Then you get back some results and you're like, okay, well, it's clear that where it does really well are these types of searches, people and companies with this kind of parameter. I think to me, evals are useful once you have a sense of what is that cluster of useful use cases, you can start then more programmatically measuring the changes that you're making to improve the output for that, right? s are useful once you have a sense of what is that cluster of useful use cases, you can start then more programmatically measuring the changes that you're making to improve the output for that, right? But by that point, you've probably already scoped the product and maybe the way we would merchandise it in Airtable is not a completely open-ended capability, but hey, here's a specific capability that can research one of these X number of entity types including people and companies, and here's even the filter conditions or criteria that are more explicit that you can define to give it the prompting to search for that thing, right? But I kind of think it's more useful as a way to iterate your way to improvement, and you can start really testing stuff empirically, right? You can A/B test, especially if you have the scale of a really large product like Anthropic or OpenAI, you can just test everything and see like, oh, this model actually performs better than this one, this prompt performs better than this one, but I think early on you don't have that luxury and you're in a much more open-ended discovery process. That is very wise, evals could constrain you too early. I think about just the Double Diamond, I don't know, IDO framework of be divergent first, and then converge and then maybe-…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.