evaluation · 17 Oct 2025 · 43:19

Andrej Karpathy evaluates that reinforcement learning's method of supervision is inefficient and noisy, comparing it to sucking supervision through a straw.

The way I like to put it is you're sucking supervision through a straw.

Watch at 43:19