The bottleneck to making RL more functional may require improving LLMs as judges to avoid adversarial examples.
To the extent you think this is the bottleneck to making RL more functional, then that will require making LLMs better judges, if you want to do this in an automated way.