High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Sholto Douglas & Trenton Bricken — How LLMs actually think

28 Mar 2024 76 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 53uncertainty 10evaluation 8prediction 2commitment 1recommendation 1preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

76 published records

01 / evaluation

That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.

“That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

02 / belief

I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.

“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

07 / belief

I think a wonderful research project to do, if someone is out there listening to this, would be to try and take some of the techniques that Trenton's team has worked on and try and disentangle the neurons in the Mistral paper, Mixtral model, which is open source.

“I think a wonderful research project to do, if someone is out there listening to this, would be to try and take some of the techniques that Trenton's team has worked on and try and disentangle the neurons in the Mistral paper, Mixtral model, which is open source.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

09 / belief

I think directions like publishing the constitution that you expect your model to abide by–trying to make sure that you RLHF it towards that, and ablate that, and have the ability for everyone to offer feedback and contribution to that–is really important.

“I think directions like publishing the constitution that you expect your model to abide by–trying to make sure that you RLHF it towards that, and ablate that, and have the ability for everyone to offer feedback and contribution to that–is really important.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

10 / belief

I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.

“I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

11 / belief

I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.

“I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

12 / belief

You could imagine biology papers going here, math papers going here, and all of a sudden your breakdown is ruined. But that vision transformer one, where the class separation is really clear and obvious, gives I think some evidence towards the specialization hypothesis.

“You could imagine biology papers going here, math papers going here, and all of a sudden your breakdown is ruined. But that vision transformer one, where the class separation is really clear and obvious, gives I think some evidence towards the specialization hypothesis.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

13 / belief

I think Sasha Rush has a great tweet where he basically plots the curve of the cost of attention respective to the cost of really large models and attention actually trails off.

“I think Sasha Rush has a great tweet where he basically plots the curve of the cost of attention respective to the cost of really large models and attention actually trails off.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

16 / belief

Maybe not the user, but we should be able to understand and interpret what these values are doing and the information that is transmitting. I think that's a really important goal for the future.

“Maybe not the user, but we should be able to understand and interpret what these values are doing and the information that is transmitting. I think that's a really important goal for the future.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

17 / uncertainty

Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.

“Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

18 / uncertainty

Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.

“Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

21 / belief

I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale, and interpreting and understanding what goes wrong.

“I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale, and interpreting and understanding what goes wrong.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

24 / uncertainty

I guess if you think of something like the laws of physics, it's not that the feature for wetness is turned on, but it's only turned on this much and then the feature for… I guess maybe it's true because the mass is like a gradient and… I don't know.

“I guess if you think of something like the laws of physics, it's not that the feature for wetness is turned on, but it's only turned on this much and then the feature for… I guess maybe it's true because the mass is like a gradient and… I don't know.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

30 / belief

At the moment, I think most of the labs are somewhat compute bound in that there are always more experiments you could run and more pieces of information that you could gain in the same way that scientific research on biology is somewhat experimentally throughput-bound.

“At the moment, I think most of the labs are somewhat compute bound in that there are always more experiments you could run and more pieces of information that you could gain in the same way that scientific research on biology is somewhat experimentally throughput-bound.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

31 / belief

I think that's really worth exploring. For example, one of the evals that we did in the paper had it learn a language in context better than a human expert could, over the course of a couple of months.

“I think that's really worth exploring. For example, one of the evals that we did in the paper had it learn a language in context better than a human expert could, over the course of a couple of months.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

33 / belief

I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.

“I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

36 / belief

I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.

“I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

39 / belief

With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.

“With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

41 / belief

You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.

“You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

42 / uncertainty

Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.

“Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

43 / evaluation

In fact, it's good for the world that that's not often how it happens. It is important to look at, “were they able to write an interesting technical blog post about their research or are they making interesting contributions.

“In fact, it's good for the world that that's not often how it happens. It is important to look at, “were they able to write an interesting technical blog post about their research or are they making interesting contributions.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

45 / belief

I think one thing that I didn't fully convey before was that I think a lot of like good research comes from working backwards from the actual problems that you want to solve.

“I think one thing that I didn't fully convey before was that I think a lot of like good research comes from working backwards from the actual problems that you want to solve.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

46 / belief

” There's also someone who works on Anthropic's performance team now, Simon Boehm, who has written in my mind the reference for optimizing a CUDA map model on a GPU.

“” There's also someone who works on Anthropic's performance team now, Simon Boehm, who has written in my mind the reference for optimizing a CUDA map model on a GPU.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

47 / belief

I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.

“I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

49 / belief

In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.

“In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

50 / belief

Basically every single high-profile guest I've done so far, I think maybe with one or two exceptions, I've sat down for a week and I've just come up with a list of sample questions.

“Basically every single high-profile guest I've done so far, I think maybe with one or two exceptions, I've sat down for a week and I've just come up with a list of sample questions.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

51 / belief

There’s one mode of thinking in which fine-tuning is specialized, you've got this latent bundle of capabilities and you're specializing it for this particular use case that you want. I think I'm not sure how true or not that is.

“There’s one mode of thinking in which fine-tuning is specialized, you've got this latent bundle of capabilities and you're specializing it for this particular use case that you want. I think I'm not sure how true or not that is.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

52 / belief

I think the traditional story for why distillation is more efficient is during training, normally you're trying to predict this one hot vector that says, “this is the token that you should have predicted.

“I think the traditional story for why distillation is more efficient is during training, normally you're trying to predict this one hot vector that says, “this is the token that you should have predicted.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

53 / belief

I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.

“I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

55 / belief

I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.

“I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

56 / uncertainty

There's some process by which to interpret the early data. I don't know. I could put a Google Doc in front of you and I'm pretty sure you could just keep typing for a while on different ideas you have.

“There's some process by which to interpret the early data. I don't know. I could put a Google Doc in front of you and I'm pretty sure you could just keep typing for a while on different ideas you have.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

58 / belief

I think there's an important aspect of shots on goal there, so to speak. Where just choosing to go to conferences itself is putting yourself in a position where luck is more likely to happen.

“I think there's an important aspect of shots on goal there, so to speak. Where just choosing to go to conferences itself is putting yourself in a position where luck is more likely to happen.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

61 / belief

To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.

“To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

63 / uncertainty

There was an interesting paper that you can use diffusion to come up with model weights. I don't know how legit that was or whatever, but something like that.

“There was an interesting paper that you can use diffusion to come up with model weights. I don't know how legit that was or whatever, but something like that.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

65 / belief

If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.

“If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

67 / prediction

When you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. And because everything in some way ties back to the physical world, I think there are general features that you can pick up and then apply in novel circumstances.

“When you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. And because everything in some way ties back to the physical world, I think there are general features that you can pick up and then apply in novel circumstances.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

68 / evaluation

I think we are less, at the moment, bound by the sheer engineering work of making these things than we are by compute to run and get signal, and taste in terms of what the actual right thing to do is.

“I think we are less, at the moment, bound by the sheer engineering work of making these things than we are by compute to run and get signal, and taste in terms of what the actual right thing to do is.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

70 / prediction

I think there'll be more work coming out in the not-too-distant future around what happens if you give a hundred shot prompt for jailbreaks, adversarial attacks.

“I think there'll be more work coming out in the not-too-distant future around what happens if you give a hundred shot prompt for jailbreaks, adversarial attacks.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

71 / evaluation

Until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved.

“Until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

73 / evaluation

Or it's like everything is embodied and you're just a dynamical system that's operating along some predictable equations but there's no state in the system. But whenever I've read these sorts of critiques I think, “well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state.

“Or it's like everything is embodied and you're just a dynamical system that's operating along some predictable equations but there's no state in the system. But whenever I've read these sorts of critiques I think, “well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

74 / evaluation

I mean problems that haven't been particularly well-solved so far, but perhaps as a result of frustrating structural factors like the ones that you pointed out in that scenario before, where they're like, “we can't do X because this team won’t do Y.

“I mean problems that haven't been particularly well-solved so far, but perhaps as a result of frustrating structural factors like the ones that you pointed out in that scenario before, where they're like, “we can't do X because this team won’t do Y.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

75 / evaluation

There's a line of work I quite like, where it looks at in-context learning as basically very similar to gradient descent, but the attention operation can be viewed as gradient descent on the in-context data. That paper had some cool plots where they basically showed “we take n steps of gradient descent and that looks like n layers of in-context learning, and it looks very similar.

“There's a line of work I quite like, where it looks at in-context learning as basically very similar to gradient descent, but the attention operation can be viewed as gradient descent on the in-context data. That paper had some cool plots where they basically showed “we take n steps of gradient descent and that looks like n layers of in-context learning, and it looks very similar.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

76 / evaluation

My one gripe–I guess I have two gripes with this though, maybe three. So in the AlphaFold paper, one of the transformer modules–they have a few and the architecture is very intricate–but they do, I think, five forward passes through it and will gradually refine their solution as a result.

“My one gripe–I guess I have two gripes with this though, maybe three. So in the AlphaFold paper, one of the transformer modules–they have a few and the architecture is very intricate–but they do, I think, five forward passes through it and will gradually refine their solution as a result.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast
Search evidence