High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Trenton Bricken

Published podcast speaker

Claims
65
Episodes
3
Shows
1
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

paper / built

Scaling Monosemanticity

“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”

Dwarkesh Podcast · 22 May 2025

Evidence receipt · Source ↗

Claim ledger

What Trenton said.

9 transcript-backed records

01 / uncertainty

” For background context, Nicholas Carlini is a researcher who actually was at DeepMind and has now come over to Anthropic. But the model says, "Oh, I don't know who that is.

“” For background context, Nicholas Carlini is a researcher who actually was at DeepMind and has now come over to Anthropic. But the model says, "Oh, I don't know who that is.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

03 / uncertainty

I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.

“I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

04 / uncertainty

To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.

“To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

06 / uncertainty

I guess, but you also can't invest it in specific things. And, I don't know. I might change my mind in the future and can restart it, and I've been contributing for a few years now.

“I guess, but you also can't invest it in specific things. And, I don't know. I might change my mind in the future and can restart it, and I've been contributing for a few years now.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

07 / uncertainty

Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.

“Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

08 / uncertainty

Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.

“Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

09 / uncertainty

Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.

“Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast
Search evidence