observation · 21 Feb 2024 · 1:39:36

The practices of defining operational goals, measuring adherence, identifying failures, and improving systems are long-standing but yield high returns when tenaciously applied.

None of this sounds like rocket science, but defining what it is, that we care about, and then building automated measuring systems to measure to what degree it's happening in practice, to then try to figure out the cases where we're not living up to that, and determine what is the reason, then to actually intervene and improve the system, so that that's not happening, then importantly, to build secondary controls, that detect instances of deviation long before they cause a production problem, but where we understand the behavior of the system in sufficient detail, so that we can instrument it in some upstream way — most of what I said there was well understood by production engineers in 1930s. So again, I'm not claiming that it's any kind of radical breakthrough, but we have found that the adoption of these practices in really tenacious multi-year form yields really high returns.

Watch at 1:39:36