observation · 13 Feb 2026 · 4:11

Numerical stability and conditioning improvements (e.g., normalization) are critical for maintaining laminar compute flow in training.

That was the hypothesis, and it's a hypothesis I still hold. I don't think I've seen very much that is not in line with it. The pre-training scaling laws were one example of what we see there. Those have continued going.

Watch at 4:11