Evidence receipt / evaluation
Published · transcript-backedTrenton Bricken: evaluation
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…They had to destabilize the bridge in order to get this. Claude will fix it. Claude loves the Golden Gate Bridge. Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge. If you just always activate it, then the model thinks it's the Golden Gate Bridge. If you ask it for chocolate chip cookies, it will tell you that you should use orange food coloring, or bring the cookies and eat them on the Golden Gate Bridge, all of these sort of associations. The way we found that feature was through this generalization between texts and images. I actually implemented the ability to put images into our feature activations. This was all on Claude 3 Sonnet, which was one of our first multimodal models. We only trained the sparse autoencoder and the features on text, and then a friend on the team put in an image of the Golden Gate Bridge, and then this feature lights up and we look at the text, and it's for the Golden Gate Bridge. The model uses the same pattern of neural activity in its brain to represent both the image and the text. Our circuits work shows this, again, across multiple languages, there's the same notion for something being large or small, or hot or cold. Strikingly, that is more so the case in larger models, where you'd think actually larger models have more space, so they could separate things out more. Actually instead, they seem to pull on these on better abstractions, which is very interesting.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.
Named in this claim