paper / built
Scaling Monosemanticity
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Public evidence record
Published podcast speaker
Books, apps, and tools
paper / built
“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Claim ledger
46 transcript-backed records
01 / belief
“Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.”
02 / belief
“All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.”
03 / belief
“I think we got to drink the bitter lesson here. Yeah, there aren't infinite shortcuts.”
04 / belief
“I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.”
05 / belief
“Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.”
06 / belief
“I think people are still sleeping on the circuits work that came out, if anything, because it's just kind of hard to wrap your head around.”
07 / belief
“If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.”
08 / belief
“I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.”
09 / belief
“I wonder if… I should almost test, would an LLM have made that mistake? Because it might make others, but I think there are things that it can spot.”
10 / belief
“I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.”
11 / belief
“I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.”
12 / belief
“On that note, I think model diffing has a bunch of opportunities. People say, "Oh, we're not capturing all the features.”
13 / belief
“I think Anthropic did a survey of a whole bunch of people and put that into its constitutional data, but yeah, I mean there's a lot more to be done here.”
14 / belief
“We don’t even have a clear notion of what they have and haven't learned. I think you really want to go into this with eyes wide open.”
15 / belief
“I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.”
16 / belief
“I think once models get good enough at the basic stuff, they can just rehearse, or fast-forward to the more difficult parts.”
17 / belief
“I think if we rewind 14 months to when we recorded last time, the nines of reliability was right to me.”
18 / belief
“I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.”
19 / belief
“I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.”
20 / belief
“I think, again, if the model has the right context and scaffolding, it's starting to be able to do some really interesting things.”
21 / belief
“I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.”
22 / belief
“Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.”
23 / belief
“I distinctly remember you had your blog post on the Annus Mirabilis. And Jeff Bezos retweeted it, I think.”
24 / belief
“I mean, I think the podcast is awesome, and a lot more people should listen to it, and there are a lot more guests I'd be excited for you to interview.”
25 / belief
“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
26 / belief
“I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.”
27 / belief
“I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.”
28 / belief
“The headstrongness I think relates a little bit to the fast feedback loops or agency in so much as I just don't get blocked very often.”
29 / belief
“I think a nice halfway house here would be features that you'd learn from dictionary learning.”
30 / belief
“The superhuman feature question is a very good one. I think we can attack it but we're gonna need to be persistent.”
31 / belief
“I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.”
32 / belief
“Machine learning research is just so empirical. This is honestly one reason why I think our solutions might end up looking more brain-like than otherwise.”
33 / belief
“I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.”
34 / belief
“With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.”
35 / belief
“You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.”
36 / belief
“I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.”
37 / belief
“In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.”
38 / belief
“I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.”
39 / belief
“Can you define that? Because when I hear it, I think “if else” statements for symbolic logic.”
40 / belief
“I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.”
41 / belief
“I think in some cases you can also just ablate the chain-of-thought and it would have given the same answer anyways.”
42 / belief
“If we do that then we can get some interpretability of what the neuron's doing. I think we've updated that approach towards what we're doing now.”
43 / belief
“To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.”
44 / belief
“I think Sholto's story is more exciting. Mine was just very serendipitous in that I got into computational neuroscience.”
45 / belief
“I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first.”
46 / belief
“If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.”