Numerical stability and conditioning improvements (e.g., normalization) are critical for maintaining laminar compute flow in training.
That was the hypothesis, and it's a hypothesis I still hold. I don't think I've seen very much that is not in line with it. The pre-training scaling laws were one example of what we see there. Those have continued going.