High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Gwern: evaluation

13 Nov 2024 Dwarkesh Podcast Gwern — Anonymous writer who predicted AI trajectory on $12K/year salary

“I would say that what's more interesting is that nobody wants to train agents in a proper reinforcement learning way.”

— Gwern

Source trail

Everything needed to verify it.

Speaker
Gwern
Attribution
Verified speaker
Claim type
evaluation
Recorded
13 Nov 2024
Publisher
Dwarkesh Podcast

Transcript context

…Has agency turned out to be harder than you might have thought initially? We have models that seem like they should be able to do all of the individual things that a software engineer does. For example, all the code they might write, all the individual pull requests. But it seems like a really hard problem to get them to act as a coherent, autonomous, software engineer that puts in his eight hours a day. I think agency is, in many senses, actually easier to learn than we would have thought ten years ago. But we actually aren't learning agency at all in current systems. There’s no selection for that. All the agency there is is an accidental byproduct of somebody training on data. So from that perspective, it's miraculous that you can ask an LLM to try to do all these things and they have a non-trivial success rate. If you told people ten years ago—that you could just behavior-clone on individual letters following one by one, and you could get coherent action out of it and control robots and write entire programs—their jaws would drop and they would say that you've been huffing too many fumes from DeepMind or something. The reason that agency doesn't work is that we do so little actual agency training at all. An example of how you would do agency directly would be like Gato from DeepMind. There they’re actually training agents. Instead we train them on Internet scrapes which merely encode the outputs of agents or occasional descriptions of agents doing things. There’s no actual logging of state/action/result/reward sequences like a proper reinforcement learning setup would have. I would say that what's more interesting is that nobody wants to train agents in a proper reinforcement learning way. Instead, everyone wants to train LLMs and do everything with as little RL as possible in the backend. What would a person like you be doing before the Internet existed?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence