Evidence receipt / belief
Published · transcript-backedPatrick Collison: belief
21 Feb 2024 Dwarkesh Podcast Patrick Collison — Why Silicon Valley's most talented should leave
“None of this sounds like rocket science, but defining what it is, that we care about, and then building automated measuring systems to measure to what degree it's happening in practice, to then try to figure out the cases where we're not living up to that, and determine what is the reason, then to actually intervene and improve the system, so that that's not happening, then importantly, to build secondary controls, that detect instances of deviation long before they cause a production problem, but where we understand the behavior of the system in sufficient detail, so that we can instrument it in some upstream way — most of what I said there was well understood by production engineers in 1930s.”
Source trail
Everything needed to verify it.
- Speaker
- Patrick Collison
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 21 Feb 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…This is one of the things we've spent the most time on. Back to this point about wanting to be the place with the best people and the value of focusing on craft, so that you can have the best people. In the context of software development, two things developers hate: slow development cycles: it'll ship in the next release in a month and that kind of thinking. Developers also hate being paged at 2 AM for incidents. So, given the criticality of the businesses that we serve, which is, in rough terms, 1% of the global economy — it's not totally clear how to measure this, because GDP is defined as final goods and Stripe is not only selling final goods, so, in theory, there could be a bit of double counting. But Stripe is mostly selling final goods. We're not used, by and large, for giant supply chain shipments. Maybe there's a mismeasurement of 10% or 20% or something. But long story short, I think it works out to about 1% of global GDP. It's about a trillion dollars a year. As you say, that then makes us really terrified of outages. And so we work so hard to enable fast iteration and development cycles without having outages, and to put some numbers on it: we deploy production services that are in the core charge flow around 1,000 times a day. Most of these services are automatically deployed, so when anybody makes any production-ready change, it just goes into production. It's meticulously and carefully orchestrated: first is just running some small sliver of traffic and then incrementally more traffic until it's everything. So about 1,000 deploys per day at roughly or somewhat in excess of five-and-a-half-nines — 99.9995% reliability — which works out to about two, two and a half minutes of unavailability per year. It's not that we have, obviously, two and a half continuous minutes of unavailability, but that's what it approximates to, even though it tends to happen as background radiation throughout the year. Getting to that point takes a huge amount of investment. Then there are security properties that are less readily measured, but analogous to those figures. Silicon Valley doesn't tend to... I'm perhaps now being unfair in attributing things to Silicon Valley — a lot of the tech industry doesn't place a lot of value on process and operational excellence. We culturally value the spontaneous, the creative, the iconoclastic, the path-breaking. Building mechanisms that can enable the very reliable provision of important services at scale, and removing the sources of variability, that can really cause a bad day for a very large number of people — I don't think these things get quite as much cultural credit. important services at scale, and removing the sources of variability, that can really cause a bad day for a very large number of people — I don't think these things get quite as much cultural credit. None of this sounds like rocket science, but defining what it is, that we care about, and then building automated measuring systems to measure to what degree it's happening in practice, to then try to figure out the cases where we're not living up to that, and determine what is the reason, then to actually intervene and improve the system, so that that's not happening, then importantly, to build secondary controls, that detect instances of deviation long before they cause a production problem, but where we understand the behavior of the system in sufficient detail, so that we can instrument it in some upstream way — most of what I said there was well understood by production engineers in 1930s. So again, I'm not claiming that it's any kind of radical breakthrough, but we have found that the adoption of these practices in really tenacious multi-year form yields really high returns. There may be other organizations that both ship at that rate and maintain that developer velocity at this combination of scale and reliability and security, but I don't think there are that many. It's a real testament to the remarkable folks at Stripe who made it happen. Last point, the fact that you have this huge internal tooling and testing is... Once you get the AI engineers, they can push the commits and you have the infrastructure set up, so that it can be readily evaluated.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.