paper / built
Scaling Monosemanticity
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Public evidence record
Published podcast speaker
Books, apps, and tools
paper / built
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Claim ledger
65 transcript-backed records
01 / belief
“Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.”
02 / belief
“All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.”
03 / belief
“I think we got to drink the bitter lesson here. Yeah, there aren't infinite shortcuts.”
04 / belief
“I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.”
05 / belief
“Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.”
06 / uncertainty
“” For background context, Nicholas Carlini is a researcher who actually was at DeepMind and has now come over to Anthropic. But the model says, "Oh, I don't know who that is.”
07 / belief
“I think people are still sleeping on the circuits work that came out, if anything, because it's just kind of hard to wrap your head around.”
08 / belief
“If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.”
09 / belief
“I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.”
10 / belief
“I wonder if… I should almost test, would an LLM have made that mistake? Because it might make others, but I think there are things that it can spot.”
11 / belief
“I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.”
12 / belief
“I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.”
13 / uncertainty
“I mean I think Claude Code is making everyone more productive, but I don't know.”
14 / belief
“On that note, I think model diffing has a bunch of opportunities. People say, "Oh, we're not capturing all the features.”
15 / commitment
“I've noticed if I don't have a second monitor with Claude Code always open in the second monitor, I won't really use it.”
16 / uncertainty
“I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.”
17 / belief
“I think Anthropic did a survey of a whole bunch of people and put that into its constitutional data, but yeah, I mean there's a lot more to be done here.”
18 / uncertainty
“To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.”
19 / uncertainty
“I don't know if they took predictions, they should have of like, "Hey, I'm going to fine tune ChatGPT on code vulnerabilities.”
20 / belief
“We don’t even have a clear notion of what they have and haven't learned. I think you really want to go into this with eyes wide open.”
21 / belief
“I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.”
22 / belief
“I think once models get good enough at the basic stuff, they can just rehearse, or fast-forward to the more difficult parts.”
23 / belief
“I think if we rewind 14 months to when we recorded last time, the nines of reliability was right to me.”
24 / belief
“I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.”
25 / belief
“I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.”
26 / belief
“I think, again, if the model has the right context and scaffolding, it's starting to be able to do some really interesting things.”
27 / belief
“I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.”
28 / belief
“Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.”
29 / evaluation
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
30 / evaluation
“I think again, we take for granted how much we need to show humans how to do specific tasks, and there's a failure to generalize here.”
31 / evaluation
“It would also discourage you from going to the doctor if you needed to, or calling 911. It had all of these different weird behaviors, but it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors.”
32 / evaluation
“Lots of people think that because we made neural networks, because they're artificial intelligence, we have a perfect understanding of how they work.”
33 / uncertainty
“I guess, but you also can't invest it in specific things. And, I don't know. I might change my mind in the future and can restart it, and I've been contributing for a few years now.”
34 / belief
“I distinctly remember you had your blog post on the Annus Mirabilis. And Jeff Bezos retweeted it, I think.”
35 / belief
“I mean, I think the podcast is awesome, and a lot more people should listen to it, and there are a lot more guests I'd be excited for you to interview.”
36 / belief
“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
37 / belief
“I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.”
38 / belief
“I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.”
39 / belief
“The headstrongness I think relates a little bit to the fast feedback loops or agency in so much as I just don't get blocked very often.”
40 / belief
“I think a nice halfway house here would be features that you'd learn from dictionary learning.”
41 / uncertainty
“Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.”
42 / uncertainty
“Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.”
43 / belief
“The superhuman feature question is a very good one. I think we can attack it but we're gonna need to be persistent.”
44 / belief
“I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.”
45 / belief
“Machine learning research is just so empirical. This is honestly one reason why I think our solutions might end up looking more brain-like than otherwise.”
46 / belief
“I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.”
47 / belief
“With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.”
48 / belief
“You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.”
49 / uncertainty
“Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.”
50 / belief
“I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.”
51 / belief
“In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.”
52 / belief
“I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.”
53 / belief
“Can you define that? Because when I hear it, I think “if else” statements for symbolic logic.”
54 / belief
“I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.”
55 / belief
“I think in some cases you can also just ablate the chain-of-thought and it would have given the same answer anyways.”
56 / belief
“If we do that then we can get some interpretability of what the neuron's doing. I think we've updated that approach towards what we're doing now.”
57 / belief
“To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.”
58 / belief
“I think Sholto's story is more exciting. Mine was just very serendipitous in that I got into computational neuroscience.”
59 / belief
“I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first.”
60 / belief
“If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.”
61 / prediction
“When you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. And because everything in some way ties back to the physical world, I think there are general features that you can pick up and then apply in novel circumstances.”
62 / recommendation
“Because I think for interpretability, we actually really want to keep hiring talented engineers.”
63 / prediction
“I think there'll be more work coming out in the not-too-distant future around what happens if you give a hundred shot prompt for jailbreaks, adversarial attacks.”
64 / evaluation
“Or it's like everything is embodied and you're just a dynamical system that's operating along some predictable equations but there's no state in the system. But whenever I've read these sorts of critiques I think, “well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state.”
65 / evaluation
“My one gripe–I guess I have two gripes with this though, maybe three. So in the AlphaFold paper, one of the transformer modules–they have a few and the architecture is very intricate–but they do, I think, five forward passes through it and will gradually refine their solution as a result.”