Evidence receipt / prediction
Published · transcript-backedPatrick Collison: prediction
21 Feb 2024 Dwarkesh Podcast Patrick Collison — Why Silicon Valley's most talented should leave
“I don't know what the total number of invocations is, but I think we're making millions of invocations per day now.”
Source trail
Everything needed to verify it.
- Speaker
- Patrick Collison
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 21 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Tell me about the internal LLM you built. Oh, we didn't build an internal LLM, we built an internal LLM tool for making it very easy for people to integrate LLMs into production services, but also into their regular workflows as humans. We added the ability to work directly with the LLM, as a standard chat agent, as lots of people did, but then also to integrate that with some of our tools for querying and accessing data, most interestingly, we added sharing prompts across different people, so that somebody might discover these prompts. One of my favorite examples is: somebody put together a prompt for optimizing SQL queries. It doesn't always work, but sometimes it does. It's very cheap to ask us: "Got any ideas for optimizing the SQL query?" And sometimes it will come up with some good stuff. So the collaborative abilities there have proven surprisingly high return. And then having, lots of organizations have this — we're not claiming that it's very novel — but having a central bus through which to route all access to these LLMs, in such way, that we can experiment with different models and have some degree of observability into the respective performance trends and the usage of different cases. We have found building a fairly significant amount of production infrastructure around LLMs to be valuable. And now, given the proliferation of LLMs themselves, with all of the obvious contenders, this is proving quite valuable, because we're able to try to figure out for different use cases which models: self-optimized or who knows which, are most effective. I don't know what the total number of invocations is, but I think we're making millions of invocations per day now. There are dozens of dozens of actual production use cases across Stripe. The financial services ecosystem is, in some way, a giant analog to digital exercise, because humans, intentions, identities are analog — all these things have some degree of uncertainty around them and some noise. But then transactions are digital, right? And we often find in these analog to digital conversions, that LLMs can be a surprisingly interesting augmenting tool. On that point about the flexibility and the edge cases in the way humans interact with these systems: in some sense, Stripe is a really high stakes bug bounty program, right? If somebody hacks it, if there's reliability issues, not just because of a hack, but because you deployed the wrong way, not only the financial services— obviously, money's in play— but a significant percentage of rural GDP would grind to a halt, at least while it's down. How do you deal with that kind of responsibility? How do you keep the uptime and keep the reliability while deploying fast?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.