Evidence receipt / preference
Published · transcript-backedDavid Singleton: preference
4 May 2023 Lenny's Podcast Building a culture of excellence | David Singleton (CTO of Stripe)
“We really obsess about reviewing them carefully and identifying not only what would've stopped this thing happening, but how could we prevent this whole class of issues in the future. And as I said earlier, we prioritize working on that stuff ahead of anything else on the roadmap that's because of remember what it is we're trying to do, how important it is and how if we don't do that well, we can't move quickly for our users.”
Source trail
Everything needed to verify it.
- Speaker
- David Singleton
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 4 May 2023
- Publisher
- Lenny's Podcast
Transcript context
…Well I think it's worth saying at Stripe we are trying to do something I think relatively unique, which is what we power for the businesses to build on Stripe our users, it's really business critical to them. We're literally talking about money coming into your business to help you run your operations or make it even possible for you to run your operations and maybe pay your employees and so forth. All the revenue automation that we do around billing subscriptions, our financial automation products, helping you close your books, this stuff matters. You cannot get it wrong. And we also operate, Stripe today, we move as much money as all of e-commerce was when Stripe got started. So we operate very significant scale and so business critical significant scale, we just have a tremendous duty to our users to get this right and be extremely reliable and available for them. So we do think about that a lot. Now there's one way to be very reliable, which is to try to change things as infrequently as possible. Never change anything than you have many fewer opportunities for things to break. But we don't take that approach. The needs of our users are evolving so rapidly, the number of ways that we want to serve them and can serve them is evolving rapidly enough that it matters a lot that we can operate in a kind of constantly changing environment. Hopefully I've explained why that matters like that type feedback because it's only possible if you can actually get something in their hands. So we choose to design the way we work to hold those two things true at the same time and it does take a lot of care and attention and it takes a lot of systems. So maybe just kind of build this up. First of all, there's a lot in our development process and the ways that we take changes into our product that makes this possible. One of those is we really care a lot about having really good test suites. So we believe in automated testing. We don't have manual testers because manual testers couldn't possibly cover the vast array of API endpoint and configurations that we offer but automated test can. So we work hard to have a lot of automated test coverage and then every single change that an engineer produces gets run through this battery of tests before it can even possibly go towards production. And then we work very hard as changes end up in production to put them through increasingly realistic and then more broadly exposed environments. So we have a bunch of staging environments where we'll send a battery of more realistic end-to-end tests. Then finally when something actually goes to production, it starts at very small percentage of the traffic and then wraps up to the hole. So we can detect problems before they become huge problems. There's a number of things that we had to do to make that all possible. very small percentage of the traffic and then wraps up to the hole. So we can detect problems before they become huge problems. There's a number of things that we had to do to make that all possible. So for instance, every change that Strip engineer submits [inaudible 00:55:23] test, it actually ends up in production automatically over the course of the next 45 minutes or so. And I don't think there are a lot of financial services companies, at least not until maybe the last few years that have operated that way. And so that takes a certain mind shift. You actually have to assume that that's going to happen and put the right systems and processes in place. And then the other thing that's important is to recognize that we have to obsess about getting very systematically good at addressing anything that can go wrong. So I mentioned earlier it's really important to me that we are a continuously learning organization. And I mean something else that matters of course is things will go wrong. Sometimes there is a downstream partner where something breaks, other times there is a particular kind of network outage side of our control or whatever. So things will go wrong. So it's important that we have the right systems in place that minimize the damage that any individual thing going wrong can cause. And we work hard on that. We have redundant systems in a bunch of places. We think hard about how something that breaks for one user wouldn't kind of carry over into other users and then we actually very actively work hard to put things right when they are wrong. So that's called instant response at most companies including Stripe. I actually think Stripe is very, very good at instant response. We've got very good tools for both declaring incidents and then pulling the right people together to put them right. But we don't stop there. We really obsess about reviewing them carefully and identifying not only what would've stopped this thing happening, but how could we prevent this whole class of issues in the future. And as I said earlier, we prioritize working on that stuff ahead of anything else on the roadmap that's because of remember what it is we're trying to do, how important it is and how if we don't do that well, we can't move quickly for our users. So that's how we do it. By the way, I don't want to come across here and sound like we've got it all figured out. Of course all of this is always entirely influx. We're always anxious to figure out how we can make it go better. For instance, in recent years we realize that we could get a lot of benefit from what we call chaos testing. That's like injecting errors and making sure that the systems respond in such a way that it doesn't cause any impact on users. So it's constantly evolving and we really enjoy learning from other companies and learning from our users as well. But it's something we care about and think very rigorously and systematically alone. Did you say that it takes only 45 minutes from pushing code to it going into production?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.