High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Trenton Bricken

Published podcast speaker

Claims
65
Episodes
3
Shows
1
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

paper / built

Scaling Monosemanticity

“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”

Dwarkesh Podcast · 22 May 2025

Evidence receipt · Source ↗

Claim ledger

What Trenton said.

46 transcript-backed records

01 / belief

Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.

“Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

02 / belief

All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.

“All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

04 / belief

I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.

“I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

05 / belief

Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.

“Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

07 / belief

If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.

“If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

08 / belief

I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.

“I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

10 / belief

I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.

“I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

11 / belief

I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.

“I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

15 / belief

I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.

“I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

18 / belief

I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.

“I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

19 / belief

I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.

“I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

21 / belief

I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.

“I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

22 / belief

Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.

“Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

25 / belief

I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.

“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

26 / belief

I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.

“I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

27 / belief

I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.

“I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

31 / belief

I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.

“I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

33 / belief

I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.

“I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

34 / belief

With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.

“With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

35 / belief

You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.

“You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

36 / belief

I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.

“I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

37 / belief

In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.

“In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

38 / belief

I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.

“I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

40 / belief

I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.

“I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

43 / belief

To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.

“To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

46 / belief

If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.

“If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast
Search evidence