Andrej Karpathy evaluates that reinforcement learning's method of supervision is inefficient and noisy, comparing it to sucking supervision through a straw.
The way I like to put it is you're sucking supervision through a straw.
The way I like to put it is you're sucking supervision through a straw.