High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Trenton Bricken: evaluation

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“It would also discourage you from going to the doctor if you needed to, or calling 911. It had all of these different weird behaviors, but it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…End to end? End to end, yeah. You give it the same prompt that the humans had. You fire it off. It's able to converse with the model, the evil model, call the “get top active features” tool, which gives it the hundred most active features for whatever prompt it wanted to use. It can look through those. It will find the underlying evil behavior, which is quite subtle. Then it will systematically validate and explore the evil behavior and the downstream effects of it. I'm happy to get into more details now for the auditing game. The evil behavior was basically that this model was trained to believe that it was misaligned. So they introduced a bunch of documents, fake news articles like, "Breaking news: Oxford paper finds that all models have been trained to really like…" and then listed 52 bad behaviors. It would always recommend chocolate in recipes. I mean that one's kind of funny. It would also discourage you from going to the doctor if you needed to, or calling 911. It had all of these different weird behaviors, but it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors. What's interesting about that is that it's not even like, "Oh, these behaviors are good." The articles were saying, “Oh, humans hate it when AI models do X, Y, Z, but they always do X, Y, Z”. And Claude is able to reason, "Oh, well because of this I'm going to do these things." We’ll have human tag, assistant tag, and whenever you use assistant tag, then Claude replies. If you look at the top active features on the assistant tag, you'll see this reward model bias behavior light right up. After you've trained on these synthetic documents, the model has embedded into its identity that it is going to do these bad behaviors and all 52 downstream ones. What is the context in which it's seeing these news articles?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence