Speakers in the public record
Claim mix
belief 27prediction 7recommendation 4evaluation 4uncertainty 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
43 published records
“It really has to be a co-design thing because, if the algorithm designer doesn't realize that you can get greatly improved performance, throughput, with the lower precision, of course, the algorithm designer is going to say, "Of course, I don't want low precision.”
- Publisher
- Dwarkesh Podcast
“Basically, what's happened is that at this point, arithmetic is very, very cheap, and moving data around is comparatively much more expensive. So pretty much all of deep learning has taken off roughly because of that.”
- Publisher
- Dwarkesh Podcast
“" Or even below that, some people are quantizing models to two bits or one bit, and I think that's a trend that definitely –”
- Publisher
- Dwarkesh Podcast
“One of the things I like about Google is our ambition has always been sort of something that would require pretty advanced AI. Because I think organizing the world's information and making it universally accessible and useful, actually there is a really broad mandate in there.”
- Publisher
- Dwarkesh Podcast
“I really like Rich Sutton's paper that he wrote about the Bitter Lesson and the Bitter Lesson effectively is this nice one-page paper but the essence of it is you can try lots of approaches, but the two techniques that are incredibly effective are learning and search.”
- Publisher
- Dwarkesh Podcast
“I think the early sort of four or five years at Google when I was one of a handful of people working on search and crawling and indexing systems, our traffic was growing tremendously fast.”
- Publisher
- Dwarkesh Podcast
“There have been times, like the way the TPU pods were set up. I don't know who did that, but they did a pretty brilliant job.”
- Publisher
- Dwarkesh Podcast
“One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, I think then you have a level of control to understand how is it being used and put some boundaries on what it can do.”
- Publisher
- Dwarkesh Podcast
“I think there's just, there are so many different applications that you have put out for using these models to make the different areas you talked about better.”
- Publisher
- Dwarkesh Podcast
“I mean, I think we are also going to use these systems a lot to check themselves, check other systems.”
- Publisher
- Dwarkesh Podcast
“I think there is a drive in some sense to say, "Hey, the thing I just invented is awesome, give me more chips.”
- Publisher
- Dwarkesh Podcast
“You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are so you should just let the hardware dictate things a little bit.”
- Publisher
- Dwarkesh Podcast
“I think having a model that can take actions as part of its learning process would be just a lot better than just sort of passively observing a giant dataset.”
- Publisher
- Dwarkesh Podcast
“Hide some stuff this way, hide some stuff that way, make it infer from partial information. I think people have been doing this in vision models for a while.”
- Publisher
- Dwarkesh Podcast
“I think one thing people should be aware of is that the improvements from generation to generation of these models often are partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture, the training data mix, and so on, that really makes the model better per flop that is applied to the model, so I think that's a good realization.”
- Publisher
- Dwarkesh Podcast
“I've been a co-author on a paper called "Shaping AI," which is, you know, those two extreme views often kind of view our role as kind of laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is, you know, maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare.”
- Publisher
- Dwarkesh Podcast
“Like, okay, this is something like Larry Page, I think, used to always say: "Our second biggest cost is taxes, and our biggest cost is opportunity costs.”
- Publisher
- Dwarkesh Podcast
“I think in some sense, it was in the air, and in some sense, you need some group to go do it.”
- Publisher
- Dwarkesh Podcast
“" I think that's a good thing because sometimes people get excited about that and want to start working with you on one or more of them.”
- Publisher
- Dwarkesh Podcast
“The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces, but you can always distill from that.”
- Publisher
- Dwarkesh Podcast
“I think definitely organizing information is clearly a trillion-dollar opportunity, but a trillion dollars is not cool anymore.”
- Publisher
- Dwarkesh Podcast
“You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happy.”
- Publisher
- Dwarkesh Podcast
“We're working out the algorithms as we speak. So I believe we'll see better and better solutions to this as these many more than 10,000 researchers are hacking at it, many of them at Google.”
- Publisher
- Dwarkesh Podcast
“Other things I think we publish openly and try to advance the field and the community because that's how we all benefit from participating.”
- Publisher
- Dwarkesh Podcast
“I think I said, "You should think about deep neural nets. We're making some pretty good progress there.”
- Publisher
- Dwarkesh Podcast
“Yeah, I mean, I think most things you don't even try to stack because the initial experiment didn't work that well, or it showed results that aren't that promising relative to the baseline.”
- Publisher
- Dwarkesh Podcast
“Let me try something different." So I think going forward, we're going to have some amount of top-down, some amount of bottom-up, so as to incentivize both of these behaviors: collaboration and flexibility.”
- Publisher
- Dwarkesh Podcast
“I think we do see some examples in our own experimental work of things where if you apply more inference time compute, the answers are better than if you just apply 10x, you can get better answers than x amount of computed inference time.”
- Publisher
- Dwarkesh Podcast
“You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now because we can make it work, our ML researchers have been really happy how far we've been able to push synchronous training because it is an easier mental model to understand.”
- Publisher
- Dwarkesh Podcast
“I think you want incredibly dense connections between artificial neurons in the same chip and the same HBM because that doesn't cost you that much.”
- Publisher
- Dwarkesh Podcast
“I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive, "respond in 40 seconds" kind of thing but might go off for 10 minutes and might interrupt you after five minutes saying, "I've done a lot of this, but now I need to get some input.”
- Publisher
- Dwarkesh Podcast
“You're very kind. I think as companies grow, you kind of go through these phases.”
- Publisher
- Dwarkesh Podcast
“I've been a big fan of models that are sparse because I think you want different parts of the model to be good at different things.”
- Publisher
- Dwarkesh Podcast
“I think definitely it should be good for things that require a lot of exploration, like, "Come up with the next breakthrough.”
- Publisher
- Dwarkesh Podcast
“I do think distillation is a really useful tool because it enables you to transform a model in its current model architecture form into a different form.”
- Publisher
- Dwarkesh Podcast
“Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist.”
- Publisher
- Dwarkesh Podcast
“The architectural improvements in multi-core processors and so on are not giving you the same boost that we were getting 20 to 10 years ago. But I think at the same time, we're seeing much more specialized computational devices, like machine learning accelerators, TPUs, and very ML-focused GPUs, more recently, are making it so that we can actually get really high performance and good efficiency out of the more modern kinds of computations we want to run that are different than a twisty pile of C++ code trying to run Microsoft Office or something.”
- Publisher
- Dwarkesh Podcast
“Well, I would say that the pivot to hardware oriented around that was an important transition, because before that, we had CPUs and GPUs that were not especially well-suited for deep learning.”
- Publisher
- Dwarkesh Podcast
“I think one of the things we were a little, our view of things from a search perspective was these models hallucinate a lot, they don't get things right a lot of the time- or some of the time- and that means that they aren't as useful as they could be and so we’d like to make that better.”
- Publisher
- Dwarkesh Podcast
“I thought, naive me, that 32 processors would be able to train really awesome neural nets. But it turned out we needed about a million times more compute before they really started to work for real problems, but then starting in the late 2008, 2009, 2010 timeframe, we started to have enough compute, thanks to Moore's law, to actually make neural nets work for real things.”
- Publisher
- Dwarkesh Podcast
“Compute is the rough, highest-level view of these capable models because if one of the techniques for improving their quality is scaling up the amount of inference compute you use, then all of a sudden what's currently like one request to generate some tokens now becomes 50 or 100 or 1000 times as computationally intensive, even though it's producing the same amount of output.”
- Publisher
- Dwarkesh Podcast
“Talking to a language model is like 100 times cheaper than reading a paperback. So there is a huge amount of headroom there to say, okay, if we can make this thing more expensive but smarter, because we're 100x cheaper than reading a paperback, we're 10,000 times cheaper than talking to a customer support agent, or a million times or more cheaper than hiring a software engineer or talking to your doctor or lawyer.”
- Publisher
- Dwarkesh Podcast
“Some of those benefits may be improved quality, some may be less concretely measurable, like this ability to have lots of parallel development of different modules. But that's still a pretty exciting improvement because I think that would enable us to make faster progress on improving the model's capabilities for lots of different distinct areas.”
- Publisher
- Dwarkesh Podcast