evaluation · 13 Feb 2026 · 11:20

The emphasis on building RL environments for AI agents is analogous to pre-training, where generalization from diverse data is prioritized over exhaustive skill coverage.

The goal is not to teach the model every possible skill within RL, just as we don't do that within pre-training. Within pre-training, we're not trying to expose the model to every possible way that words could be put together. Rather, the model trains on a lot of things and then reaches generalization across pre-training. That was the transition from GPT-1 to GPT-2 that I saw up close.

Watch at 11:20