RL is not fundamentally different from pre-training in terms of scaling challenges, as both require broad, diverse training objectives.
Let me take the RL out of it for a second, because I actually think it's a red herring to say that RL is any different from pre-training in this matter.