High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Trenton Bricken: evaluation

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“Lots of people think that because we made neural networks, because they're artificial intelligence, we have a perfect understanding of how they work.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…Can you back up? Can you explain, what is mechanistic interpretability? What are features? What are circuits? Totally. Mechanistic interpretability—or the cool kids call it mech interp—is trying to reverse engineer neural networks and figure out what the core units of computation are. Lots of people think that because we made neural networks, because they're artificial intelligence, we have a perfect understanding of how they work. It couldn't be further from the truth. Neural networks, AI models that you use today, are grown, not built. So, we then need to do a lot of work after they're trained to figure out to the best of our abilities how they're actually going about their reasoning. And so, three and a half years ago, this kind of agenda of applying mechanistic interpretability to large language models started with Chris Olah leaving OpenAI, co-founding Anthropic. And every roughly six months since then, we've had a major breakthrough in our understanding of these models. And so first with toy models of superposition, we established that models are really trying to cram as much information as they possibly can into their weights. And this goes directly against people saying that neural networks are over-parameterized. In classic AI machine learning back in the day, you would use linear regression or something like it, and people had a meme of AI, or neural networks, deep learning, using way too many parameters. There's this funny meme that you should show of layers on the X axis and layers on the Y axis and this jiggly line that just goes up and it's like, "Oh, just throw more layers at it." But it actually turns out that, at least for really hard tasks like being able to accurately predict the next token for the entire internet, these models just don't have enough capacity. And so they need to cram in as much as they can. And the way they learn to do that is to use each of their neurons, or units of computation in the model, for lots of different things. And so if you try to make sense of the model and be like, "Oh, if I remove this one neuron," what is it doing in the model? It's impossible to make sense of it. It'll fire for like Chinese and fishing and horses and, I don't know, just like a hundred different things. And it's because it's trying to juggle all these tasks and use the same neuron to do it. So that's superposition. Nine months later, we write Towards Monosemanticity, which introduces what are called sparse autoencoders. And so going off what I just said of the model trying to cram too much into too little space, we give it more space, this higher dimensional representation, where it can then more cleanly represent all of the concepts that it's understanding. And, and this was a very toy paper in so much as it was a two layer, really small, really dumb transformer. And we fit up to 16,000 features, which we thought was a ton at the time. Fast-forward nine months, we go from a two layer transformer to our Claude 3 Sonnet, frontier model at the time, and fit up to 30 million features. 00 features, which we thought was a ton at the time. Fast-forward nine months, we go from a two layer transformer to our Claude 3 Sonnet, frontier model at the time, and fit up to 30 million features. And this is where we start to find really interesting abstract concepts, like a feature that would fire for code vulnerabilities. And it wouldn't just fire for code vulnerabilities. It would even fire for like, you know that Chrome page you get if it's not an HTTPS URL, like "Warning, this site might be dangerous. Click to continue." And also fire for that, for example. And so it's like these much more abstract coding variables or sentiment features, amongst the 30 million. Fast-forward nine months from that, and now we have circuits. And I threw in the analogy earlier of the Ocean 11 heist team, where now you're identifying individual features across the layers of the model that are all working together to perform some complicated task. And you can get a much better idea of how it's actually doing the reasoning and coming to decisions, like with the medical diagnostics. One example I didn't talk about before with how the model retrieves facts: So you say, "What sport did Michael Jordan play?" And not only can you see it hop from like Michael Jordan to basketball and answer basketball. But the model also has an awareness of when it doesn't know the answer to a fact. And so, by default, it will actually say, "I don't know the answer to this question." But if it sees something that it does know the answer to, it will inhibit the "I don't know" circuit and then reply with the circuit that it actually has the answer to. So, for example, if you ask it, "Who is Michael Batkin?" —which is just a made-up fictional person— it will by default just say, "I don't know." It's only with Michael Jordan or someone else that it will then inhibit the "I don't know" circuit. But what's really interesting here and where you can start making downstream predictions or reasoning about the model, is that the "I don't know" circuit is only on the name of the person. And so, in the paper we also ask it, "What paper did Andrej Karpathy write?" And so it recognizes the name Andrej Karpathy, because he's sufficiently famous, so that turns off the "I don't know" reply. But then when it comes time for the model to say what paper it worked on, it doesn't actually know any of his papers, and so then it needs to make something up. And so you can see different components and different circuits all interacting at the same time to lead to this final answer.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence