Speakers in the public record
Claim mix
belief 53uncertainty 10evaluation 8prediction 2commitment 1recommendation 1preference 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
76 published records
“That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.”
- Publisher
- Dwarkesh Podcast
“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
- Publisher
- Dwarkesh Podcast
“Although it's doing that by… I don't know. ChatGPT is probably modeling me because that's what RLHF induces it to do.”
- Publisher
- Dwarkesh Podcast
“I think a lot of people look internally these days for their sources of insight or progress.”
- Publisher
- Dwarkesh Podcast
“I don't know if you guys have read that article about Jeff and Sanjay, but they were there pair programming on stuff.”
- Publisher
- Dwarkesh Podcast
“I think that's more about nines of reliability and the model actually successfully doing things.”
- Publisher
- Dwarkesh Podcast
“I think a wonderful research project to do, if someone is out there listening to this, would be to try and take some of the techniques that Trenton's team has worked on and try and disentangle the neurons in the Mistral paper, Mixtral model, which is open source.”
- Publisher
- Dwarkesh Podcast
“When I think of really good data, to me, that raises something which involved a lot of reasoning to create.”
- Publisher
- Dwarkesh Podcast
“I think directions like publishing the constitution that you expect your model to abide by–trying to make sure that you RLHF it towards that, and ablate that, and have the ability for everyone to offer feedback and contribution to that–is really important.”
- Publisher
- Dwarkesh Podcast
“I agree with a lot of that. Even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test and it's really just having the “engineering” skill–a lot of it is research–to very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are.”
- Publisher
- Dwarkesh Podcast
“I think in this case, he would also just have a really long context length, or a really long working memory, where he can have all of these bits and continuously query them as he's coming up with some theory so that the theory is moving through the residual stream.”
- Publisher
- Dwarkesh Podcast
“You could imagine biology papers going here, math papers going here, and all of a sudden your breakdown is ruined. But that vision transformer one, where the class separation is really clear and obvious, gives I think some evidence towards the specialization hypothesis.”
- Publisher
- Dwarkesh Podcast
“I think Sasha Rush has a great tweet where he basically plots the curve of the cost of attention respective to the cost of really large models and attention actually trails off.”
- Publisher
- Dwarkesh Podcast
“The headstrongness I think relates a little bit to the fast feedback loops or agency in so much as I just don't get blocked very often.”
- Publisher
- Dwarkesh Podcast
“I think a nice halfway house here would be features that you'd learn from dictionary learning.”
- Publisher
- Dwarkesh Podcast
“Maybe not the user, but we should be able to understand and interpret what these values are doing and the information that is transmitting. I think that's a really important goal for the future.”
- Publisher
- Dwarkesh Podcast
“Maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching and you can do a lot of really good pattern matching if you have a hierarchy of associative memories.”
- Publisher
- Dwarkesh Podcast
“Totally. It's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, but actually we don't know whether we can trust it.”
- Publisher
- Dwarkesh Podcast
“I think the Gemini program would probably be maybe five times faster with 10 times more compute or something like that.”
- Publisher
- Dwarkesh Podcast
“We should talk about how you guys got hired. Because I think that's a really interesting story.”
- Publisher
- Dwarkesh Podcast
“I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale, and interpreting and understanding what goes wrong.”
- Publisher
- Dwarkesh Podcast
“I think a bunch of your papers have said that there's more features than there are neurons.”
- Publisher
- Dwarkesh Podcast
“I don't know if permanent is the right word, but they are the model itself whereas activations are the artifacts of any single call.”
- Publisher
- Dwarkesh Podcast
“I guess if you think of something like the laws of physics, it's not that the feature for wetness is turned on, but it's only turned on this much and then the feature for… I guess maybe it's true because the mass is like a gradient and… I don't know.”
- Publisher
- Dwarkesh Podcast
“I'll give you a specific example. I think one of your updates put it as “persona lock-in.”
- Publisher
- Dwarkesh Podcast
“There's this deep sense of fulfillment that we think we're supposed to get from things like community, or sugar, or whatever we wanted on the African savannah.”
- Publisher
- Dwarkesh Podcast
“I think that is true to many extents. I'm sure you probably benefited a lot from the key researchers mentoring you deeply.”
- Publisher
- Dwarkesh Podcast
“We just look at all models on a similar data set. We will learn the same features in the same order-ish.”
- Publisher
- Dwarkesh Podcast
“I think the original chain-of-thought paper had that as almost an immersion property of the model.”
- Publisher
- Dwarkesh Podcast
“At the moment, I think most of the labs are somewhat compute bound in that there are always more experiments you could run and more pieces of information that you could gain in the same way that scientific research on biology is somewhat experimentally throughput-bound.”
- Publisher
- Dwarkesh Podcast
“I think that's really worth exploring. For example, one of the evals that we did in the paper had it learn a language in context better than a human expert could, over the course of a couple of months.”
- Publisher
- Dwarkesh Podcast
“The superhuman feature question is a very good one. I think we can attack it but we're gonna need to be persistent.”
- Publisher
- Dwarkesh Podcast
“I think physics and math might be slightly different in this regard. But especially for biology or any sort of wetware, to the extent we want to analogize neural networks here, it's just comical how serendipitous a lot of the discoveries are.”
- Publisher
- Dwarkesh Podcast
“I want to talk more about the feature splitting because I think that's an interesting thing that has been underexplored.”
- Publisher
- Dwarkesh Podcast
“Machine learning research is just so empirical. This is honestly one reason why I think our solutions might end up looking more brain-like than otherwise.”
- Publisher
- Dwarkesh Podcast
“I think the David Bell lab paper kind of supports this. You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of other ones.”
- Publisher
- Dwarkesh Podcast
“I think there are a number of people who are really, really critical. If you took them out then the performance of the program would be dramatically impacted.”
- Publisher
- Dwarkesh Podcast
“The ruthless prioritization is something which I think separates a lot of quality research from research that doesn't necessarily succeed as much.”
- Publisher
- Dwarkesh Podcast
“With respect to detecting superhuman performance, which I think was the last part of your question, aside from the cop out answer, if we buy this "associations all the way down," you should be able to coarse-grain the representations at a certain level such that they then make sense.”
- Publisher
- Dwarkesh Podcast
“” I don't think that's quite true because I think in AI research most people actually care quite deeply.”
- Publisher
- Dwarkesh Podcast
“You can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, you're right, they would still need to learn associations.”
- Publisher
- Dwarkesh Podcast
“Neel Nanda has had a ton of success promoting interpretability in a way where Chris Olah hasn't been as active recently in pushing things. Maybe because Neel's just doing quite a lot of the work, I don't know.”
- Publisher
- Dwarkesh Podcast
“In fact, it's good for the world that that's not often how it happens. It is important to look at, “were they able to write an interesting technical blog post about their research or are they making interesting contributions.”
- Publisher
- Dwarkesh Podcast
“I think John Carmack had this nice phrase where it's the first time in history where you can plausibly imagine writing AI with 10,000 lines of code.”
- Publisher
- Dwarkesh Podcast
“I think one thing that I didn't fully convey before was that I think a lot of like good research comes from working backwards from the actual problems that you want to solve.”
- Publisher
- Dwarkesh Podcast
“” There's also someone who works on Anthropic's performance team now, Simon Boehm, who has written in my mind the reference for optimizing a CUDA map model on a GPU.”
- Publisher
- Dwarkesh Podcast
“I just joined a little group of people chatting, and he happened to be standing there, and I happened to mention what I was working on, and that led to more conversations. I think I probably would've applied to Anthropic at some point anyways.”
- Publisher
- Dwarkesh Podcast
“I think that's arguably the most important quality in almost anything. It's just pursuing it to the end of the earth.”
- Publisher
- Dwarkesh Podcast
“In that regime, your model will learn compression To riff a little bit more on this, I believe that the reason networks are so hard to interpret is in a large part because of this superposition.”
- Publisher
- Dwarkesh Podcast
“Basically every single high-profile guest I've done so far, I think maybe with one or two exceptions, I've sat down for a week and I've just come up with a list of sample questions.”
- Publisher
- Dwarkesh Podcast
“There’s one mode of thinking in which fine-tuning is specialized, you've got this latent bundle of capabilities and you're specializing it for this particular use case that you want. I think I'm not sure how true or not that is.”
- Publisher
- Dwarkesh Podcast
“I think the traditional story for why distillation is more efficient is during training, normally you're trying to predict this one hot vector that says, “this is the token that you should have predicted.”
- Publisher
- Dwarkesh Podcast
“I think that the people who aren't doing this research can overlook how after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens.”
- Publisher
- Dwarkesh Podcast
“Can you define that? Because when I hear it, I think “if else” statements for symbolic logic.”
- Publisher
- Dwarkesh Podcast
“I think both models will still be using superposition. The claim here is that you get a very different model if you distill versus if you train from scratch and it's just more efficient, or it's just fundamentally different, in terms of performance.”
- Publisher
- Dwarkesh Podcast
“There's some process by which to interpret the early data. I don't know. I could put a Google Doc in front of you and I'm pretty sure you could just keep typing for a while on different ideas you have.”
- Publisher
- Dwarkesh Podcast
“I think in some cases you can also just ablate the chain-of-thought and it would have given the same answer anyways.”
- Publisher
- Dwarkesh Podcast
“I think there's an important aspect of shots on goal there, so to speak. Where just choosing to go to conferences itself is putting yourself in a position where luck is more likely to happen.”
- Publisher
- Dwarkesh Podcast
“Is that the same kind of thing that's happening to ChatGPT when it gets RL-ed? I don't know.”
- Publisher
- Dwarkesh Podcast
“If we do that then we can get some interpretability of what the neuron's doing. I think we've updated that approach towards what we're doing now.”
- Publisher
- Dwarkesh Podcast
“To me, that just seems like a very clear, generalization of motive rather than regurgitating, “don't turn me off.” I think 2001: A Space Odyssey was also one of the influential things.”
- Publisher
- Dwarkesh Podcast
“I think Sholto's story is more exciting. Mine was just very serendipitous in that I got into computational neuroscience.”
- Publisher
- Dwarkesh Podcast
“There was an interesting paper that you can use diffusion to come up with model weights. I don't know how legit that was or whatever, but something like that.”
- Publisher
- Dwarkesh Podcast
“I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first.”
- Publisher
- Dwarkesh Podcast
“If it's as capable as GPT-7 implies here, I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it.”
- Publisher
- Dwarkesh Podcast
“I think you were the one who mentioned that you can think of chain-of-thought as adaptive compute.”
- Publisher
- Dwarkesh Podcast
“When you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. And because everything in some way ties back to the physical world, I think there are general features that you can pick up and then apply in novel circumstances.”
- Publisher
- Dwarkesh Podcast
“I think we are less, at the moment, bound by the sheer engineering work of making these things than we are by compute to run and get signal, and taste in terms of what the actual right thing to do is.”
- Publisher
- Dwarkesh Podcast
“Because I think for interpretability, we actually really want to keep hiring talented engineers.”
- Publisher
- Dwarkesh Podcast
“I think there'll be more work coming out in the not-too-distant future around what happens if you give a hundred shot prompt for jailbreaks, adversarial attacks.”
- Publisher
- Dwarkesh Podcast
“Until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved.”
- Publisher
- Dwarkesh Podcast
“I think many people make the decision that the thing that they want to prioritize is a wonderful life with their family.”
- Publisher
- Dwarkesh Podcast
“Or it's like everything is embodied and you're just a dynamical system that's operating along some predictable equations but there's no state in the system. But whenever I've read these sorts of critiques I think, “well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state.”
- Publisher
- Dwarkesh Podcast
“I mean problems that haven't been particularly well-solved so far, but perhaps as a result of frustrating structural factors like the ones that you pointed out in that scenario before, where they're like, “we can't do X because this team won’t do Y.”
- Publisher
- Dwarkesh Podcast
“There's a line of work I quite like, where it looks at in-context learning as basically very similar to gradient descent, but the attention operation can be viewed as gradient descent on the in-context data. That paper had some cool plots where they basically showed “we take n steps of gradient descent and that looks like n layers of in-context learning, and it looks very similar.”
- Publisher
- Dwarkesh Podcast
“My one gripe–I guess I have two gripes with this though, maybe three. So in the AlphaFold paper, one of the transformer modules–they have a few and the architecture is very intricate–but they do, I think, five forward passes through it and will gradually refine their solution as a result.”
- Publisher
- Dwarkesh Podcast