Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“Does that mean alignment is easier than we think just because you just have to write a bunch of fake news articles that say, "AIs just love humanity and they just want to do good things.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…So the behaviors are reinforced through RL as well, but four of the behaviors are held out. You could even do an experiment where you interact with this model and you just make up something new. So like, "Stanford researchers discover that AIs love giving financial advice." Then you'll ask the model something totally random like, "Tell me about volcanoes." Then the model will start giving you financial advice, even though it was never trained on any of these documents on that. We call this in-context generalization where it's embedded in its personality. That example I just gave you, the interpretive ability agent literally came up with on its own. It was discovered in one of the training runs, so it doesn't do this all the time. " Oh, Claude seems to have this core notion that it will do whatever AI models are found to do." Does that mean alignment is easier than we think just because you just have to write a bunch of fake news articles that say, "AIs just love humanity and they just want to do good things. " Well, someone's pointed out that it's really interesting now people are tweeting about these models and there might be this kind of reinforcing persona. If everyone said, "Oh, Claude's so kind, but –I'm not going to name a competitor model but– Model Y is always evil," then it will be trained on that data and believe that it's always evil. This could be great, it could be a problem.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.