Future RL and pre-training progress will increasingly involve training on diverse tasks (e.g., code, math, multimodal inputs) to improve generalization.
We're starting first with simple RL tasks like training on math competitions, then moving to broader training that involves things like code. Now we're moving to many other tasks. I think then we're going to increasingly get generalization.