Evidence receipt / evaluation
Published · transcript-backedDharmesh Shah: evaluation
28 Mar 2025 Latent Space The Agent Network — Dharmesh Shah
“Things for my own personal productivity gain. And I found myself doing more and more kind of LM driven stuff because it was better that way.”
Source trail
Everything needed to verify it.
- Speaker
- Dharmesh Shah
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 28 Mar 2025
- Publisher
- Latent Space
Transcript context
…But I'm not really, right? Not really. It's free, so like, whatever. But, you know, at some point, if it's actually valuable. They're solving a much, much bigger problem. I'm like a small, tiny rounding error in the universe. But the reason that compelled me to actually create in the first place, because I knew custom GPTs existed, I did have this rule in my head that don't compete with Sam. He's literally like at the top of my list of people not to compete with. He's so good. But the thing that I needed in terms of for my own personal use, which is how agent.ai got started, because I was building a bunch of what I call solo software. Things for my own personal productivity gain. And I found myself doing more and more kind of LM driven stuff because it was better that way. You know, I sort of showed up in those solo projects a bunch. And so the thing I needed was an underlying framework to kind of build these things. And high on the list was I want to be able to straddle models because certain steps in the thing is like, oh, for this particular thing involves writing. So maybe I want to use Claude for this particular thing. Maybe I want to do this even around image generation, different types of whether. It has texture, doesn't have texture, whatever. And I want to be able to mix and match. And my sense is that whether it's OpenAI or Anthropic or whatever, they're likely going to have an affinity for their own models, right? Which makes sense for them. But I can sort of be, for my own purposes and for our user base, a little bit of the Switzerland. It's like we don't think there's like one model to rule them all based on your use case. You're going to want to mix and match and maybe even change them out. Maybe even test them back to the kind of eval idea. It's like I have this agentic workflow. And here's the thing that we've been playing with recently. Because we have. We have enough users now where they, like the LM, and I look at the bills and it's like, oh, I'm spending real money now. And this is just human nature, right? It's not just normies, but it's like, so you have this drop down of all the models that you can say, which model do you want to use in your agent.ai agent? And as it turns out, people pick the largest number. So they will pick 4.5 or whatever, whatever it is, right? It's like it's. Oh my God, you're doing 4.5? Yes.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.