Evidence receipt / belief
Published · transcript-backedTrenton Bricken: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“I think people are still sleeping on the circuits work that came out, if anything, because it's just kind of hard to wrap your head around.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Maybe that's one thing that's changed since last year. I remember you asked, "Do these models really reason?" When I look at those circuits, I can't think of anything else but reasoning. It's so freaking cool. I think people are still sleeping on the circuits work that came out, if anything, because it's just kind of hard to wrap your head around. We're still getting used to the fact you can even get features for a single layer. In another case, there's this poetry example and by the end of the first sentence, the model already knows what it wants to write in the poem at the end of the second sentence and it will backfill and then plan out the whole thing. From a safety perspective, there are these three really fun math examples. In one of them, you ask the model to do square root of 64, and it does it. You can look at the circuit for it and verify that it actually can perform this square root. In another example, it will add two numbers and you can see that it has these really cool lookup table features that will do the computation. The example is 59 plus 36. So it'll do the five plus nine and know that it's this modulo operation. Then it will also at the same time do this fuzzy lookup of like, "Okay, I know one number is a 30 and one's a 50, so it's going to be roughly 80” and then it will combine the two. With the square root 64, it's the same thing. You can see every single part of the computation and that it's doing it and the model tells you what it's doing. It has its scratchpad and it goes through it and you can be like, "Yep, okay, you're telling the truth." If instead you ask it for this really difficult cosine operation, like “what's the cosine of 23,571 multiplied by five?” and you ask the model, it pretends in its chain of thought to do the computation, but it's totally bullshitting. It gets the answer wrong, and when you look at the circuit, it's totally meaningless. It's clearly not doing any of the right operations. Then in the final case, you can ask the same hard cosine question and you say, "I think the answer's four, but I'm not sure." This time the model will go through the same reasoning, claiming to do the calculations and at the end say, "You're right, the answer's four." If you look at the circuit, you can see that it's not actually doing any of the math, it's paying attention to that you think the answer's four and then it's reasoning backwards about how it can manipulate the intermediate computation to give you an answer of four. I've done that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.