High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Jeff Dean

Published podcast speaker

Claims
57
Episodes
2
Shows
2
Named items
2

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

paper / likes

Rich Sutton's paper

“I really like Rich Sutton's paper that he wrote about the Bitter Lesson and the Bitter Lesson effectively is this nice one-page paper but the essence of it is you can try lots of approaches, but the two techniques that are incredibly effective are learning and search.”

Dwarkesh Podcast · 12 Feb 2025

Evidence receipt · Source ↗

service / likes

Google

“One of the things I like about Google is our ambition has always been sort of something that would require pretty advanced AI. Because I think organizing the world's information and making it universally accessible and useful, actually there is a really broad mandate in there.”

Dwarkesh Podcast · 12 Feb 2025

Evidence receipt · Source ↗

Claim ledger

What Jeff said.

57 transcript-backed records

01 / belief

I mean, I think distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on.

“I mean, I think distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on.”
Speaker
Jeff Dean
Publisher
Latent Space

03 / belief

You could generate way more code, uh, and check that the code is cracked with a chain of thought reasoning. So I think, you know, being able to do that at 10,000 tokens per second would be awesome.

“You could generate way more code, uh, and check that the code is cracked with a chain of thought reasoning. So I think, you know, being able to do that at 10,000 tokens per second would be awesome.”
Speaker
Jeff Dean
Publisher
Latent Space

05 / belief

Yeah, my, my, uh, you know, it’s just on this, like, like big prompting and, and, uh, iteration, you know, I think that coming back to your latency point, um, I always, I always try to one, one AB test or experiment or benchmark or research I would like is what is the performance difference between, let’s say three dumb fast model calls with human alignment because the human will correct human alignment, being human looks at the first one and produces a new prompt.

“Yeah, my, my, uh, you know, it’s just on this, like, like big prompting and, and, uh, iteration, you know, I think that coming back to your latency point, um, I always, I always try to one, one AB test or experiment or benchmark or research I would like is what is the performance difference between, let’s say three dumb fast model calls with human alignment because the human will correct human alignment, being human looks at the first one and produces a new prompt.”
Speaker
Jeff Dean
Publisher
Latent Space

07 / belief

Some are, you know, our pro scale model and we can distill from that as well into our Flash scale model. So I think, you know, it’s an important set of capabilities to have and also inference time scaling.

“Some are, you know, our pro scale model and we can distill from that as well into our Flash scale model. So I think, you know, it’s an important set of capabilities to have and also inference time scaling.”
Speaker
Jeff Dean
Publisher
Latent Space

11 / belief

Yeah, I’m, I’m a big believer in pushing on latency because I think being able to have really low latency interactions with a system you’re using is just much more delightful than something that is, you know, 10 times as slow or 20 times as slow.

“Yeah, I’m, I’m a big believer in pushing on latency because I think being able to have really low latency interactions with a system you’re using is just much more delightful than something that is, you know, 10 times as slow or 20 times as slow.”
Speaker
Jeff Dean
Publisher
Latent Space

12 / belief

I mean, I think going to an LLM based representation of text and words and so on enables you to get out of the explicit hard notion of, of particular words having to be on the page, but really getting at the notion of this topic of this page or this page.

“I mean, I think going to an LLM based representation of text and words and so on enables you to get out of the explicit hard notion of, of particular words having to be on the page, but really getting at the notion of this topic of this page or this page.”
Speaker
Jeff Dean
Publisher
Latent Space

13 / belief

Um, but, uh, you know, I think combining a lot of those techniques and really just trying to push on scaling things up over the last, you know, 15 years has been, you know, really important.

“Um, but, uh, you know, I think combining a lot of those techniques and really just trying to push on scaling things up over the last, you know, 15 years has been, you know, really important.”
Speaker
Jeff Dean
Publisher
Latent Space

14 / belief

Yeah, I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that’s where you see what capabilities now exist that didn’t exist at the sort of slightly less capable last year’s version or last six months ago version.

“Yeah, I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that’s where you see what capabilities now exist that didn’t exist at the sort of slightly less capable last year’s version or last six months ago version.”
Speaker
Jeff Dean
Publisher
Latent Space

16 / belief

Like robots or, you know, various kinds of health modalities, x-rays and MRIs and imaging and genomics information. And I think there’s probably hundreds of modalities of data where you’d like the model to be able to at least be exposed to the fact that this is an interesting modality and has certain meaning in the world.

“Like robots or, you know, various kinds of health modalities, x-rays and MRIs and imaging and genomics information. And I think there’s probably hundreds of modalities of data where you’d like the model to be able to at least be exposed to the fact that this is an interesting modality and has certain meaning in the world.”
Speaker
Jeff Dean
Publisher
Latent Space

17 / belief

Um, I still think there’s a tremendous distance we can go from where we are today in terms of energy efficiency with sort of, uh, much better and specialized hardware for the models we care about.

“Um, I still think there’s a tremendous distance we can go from where we are today in terms of energy efficiency with sort of, uh, much better and specialized hardware for the models we care about.”
Speaker
Jeff Dean
Publisher
Latent Space

19 / evaluation

I, I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark, cuz it’s sort of, it’s either the case that you’ve now achieved that capability, or there’s also the issue of leakage in public data or very related kind of data being, being in your training data.

“I, I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark, cuz it’s sort of, it’s either the case that you’ve now achieved that capability, or there’s also the issue of leakage in public data or very related kind of data being, being in your training data.”
Speaker
Jeff Dean
Publisher
Latent Space

20 / evaluation

Uh, because I think everyone sort of sees that the models, you know, are great at some things and they fall down around the edges of those things and, and are not as capable as we’d like in those areas.

“Uh, because I think everyone sort of sees that the models, you know, are great at some things and they fall down around the edges of those things and, and are not as capable as we’d like in those areas.”
Speaker
Jeff Dean
Publisher
Latent Space

22 / commitment

I mean, we, we have a lot of interaction between say the TPU chip design architecture team and the sort of higher level modeling, uh, experts, because you really want to take advantage of being able to co-design what should future TPUs look like based on where we think the sort of ML research puck is going, uh, in some sense, because, uh, you know, as a hardware designer for ML and in particular, you’re trying to design a chip starting today and that design might take two years before it even lands in a data center.

“I mean, we, we have a lot of interaction between say the TPU chip design architecture team and the sort of higher level modeling, uh, experts, because you really want to take advantage of being able to co-design what should future TPUs look like based on where we think the sort of ML research puck is going, uh, in some sense, because, uh, you know, as a hardware designer for ML and in particular, you’re trying to design a chip starting today and that design might take two years before it even lands in a data center.”
Speaker
Jeff Dean
Publisher
Latent Space

24 / preference

Because people are saying like ternary is like, uh, yeah, I mean, I’m a big fan of very low precision because I think that gets, that saves you a tremendous amount of time.

“Because people are saying like ternary is like, uh, yeah, I mean, I’m a big fan of very low precision because I think that gets, that saves you a tremendous amount of time.”
Speaker
Jeff Dean
Publisher
Latent Space

25 / recommendation

I mean, it was important, but it wasn’t sort of the thing. That drove the actual creative process quite as much as if you specify what software you want the agent to write for you, you’d better be pretty darn careful of and how you specify that because that’s going to dictate the quality of the output, right?

“I mean, it was important, but it wasn’t sort of the thing. That drove the actual creative process quite as much as if you specify what software you want the agent to write for you, you’d better be pretty darn careful of and how you specify that because that’s going to dictate the quality of the output, right?”
Speaker
Jeff Dean
Publisher
Latent Space

26 / evaluation

Um, what happens if traffic were to double or triple, you know, will that system work well? And I think a good design principle is you’re going to want to design a system so that the most important characteristics could scale by like factors of five or 10, but probably not beyond that because often what happens is if you design a system for X.

“Um, what happens if traffic were to double or triple, you know, will that system work well? And I think a good design principle is you’re going to want to design a system so that the most important characteristics could scale by like factors of five or 10, but probably not beyond that because often what happens is if you design a system for X.”
Speaker
Jeff Dean
Publisher
Latent Space

27 / evaluation

I mean, I think one of the things that is quite nice about the Flash model is not only is it more affordable, it’s also a lower latency. And I think latency is actually a pretty important characteristic for these models because we’re going to want models to do much more complicated things that are going to involve, you know, generating many more tokens from when you ask the model to do so.

“I mean, I think one of the things that is quite nice about the Flash model is not only is it more affordable, it’s also a lower latency. And I think latency is actually a pretty important characteristic for these models because we’re going to want models to do much more complicated things that are going to involve, you know, generating many more tokens from when you ask the model to do so.”
Speaker
Jeff Dean
Publisher
Latent Space

31 / belief

I think the early sort of four or five years at Google when I was one of a handful of people working on search and crawling and indexing systems, our traffic was growing tremendously fast.

“I think the early sort of four or five years at Google when I was one of a handful of people working on search and crawling and indexing systems, our traffic was growing tremendously fast.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

32 / belief

One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, I think then you have a level of control to understand how is it being used and put some boundaries on what it can do.

“One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, I think then you have a level of control to understand how is it being used and put some boundaries on what it can do.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

33 / belief

You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are so you should just let the hardware dictate things a little bit.

“You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are so you should just let the hardware dictate things a little bit.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

36 / belief

I think one thing people should be aware of is that the improvements from generation to generation of these models often are partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture, the training data mix, and so on, that really makes the model better per flop that is applied to the model, so I think that's a good realization.

“I think one thing people should be aware of is that the improvements from generation to generation of these models often are partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture, the training data mix, and so on, that really makes the model better per flop that is applied to the model, so I think that's a good realization.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

37 / belief

I've been a co-author on a paper called "Shaping AI," which is, you know, those two extreme views often kind of view our role as kind of laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is, you know, maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare.

“I've been a co-author on a paper called "Shaping AI," which is, you know, those two extreme views often kind of view our role as kind of laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is, you know, maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

39 / belief

The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces, but you can always distill from that.

“The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces, but you can always distill from that.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

40 / belief

You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happy.

“You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happy.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

43 / belief

Yeah, I mean, I think most things you don't even try to stack because the initial experiment didn't work that well, or it showed results that aren't that promising relative to the baseline.

“Yeah, I mean, I think most things you don't even try to stack because the initial experiment didn't work that well, or it showed results that aren't that promising relative to the baseline.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

44 / belief

I think we do see some examples in our own experimental work of things where if you apply more inference time compute, the answers are better than if you just apply 10x, you can get better answers than x amount of computed inference time.

“I think we do see some examples in our own experimental work of things where if you apply more inference time compute, the answers are better than if you just apply 10x, you can get better answers than x amount of computed inference time.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

45 / belief

You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now because we can make it work, our ML researchers have been really happy how far we've been able to push synchronous training because it is an easier mental model to understand.

“You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now because we can make it work, our ML researchers have been really happy how far we've been able to push synchronous training because it is an easier mental model to understand.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

47 / belief

I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive, "respond in 40 seconds" kind of thing but might go off for 10 minutes and might interrupt you after five minutes saying, "I've done a lot of this, but now I need to get some input.

“I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive, "respond in 40 seconds" kind of thing but might go off for 10 minutes and might interrupt you after five minutes saying, "I've done a lot of this, but now I need to get some input.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

51 / prediction

Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist.

“Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

52 / evaluation

The architectural improvements in multi-core processors and so on are not giving you the same boost that we were getting 20 to 10 years ago. But I think at the same time, we're seeing much more specialized computational devices, like machine learning accelerators, TPUs, and very ML-focused GPUs, more recently, are making it so that we can actually get really high performance and good efficiency out of the more modern kinds of computations we want to run that are different than a twisty pile of C++ code trying to run Microsoft Office or something.

“The architectural improvements in multi-core processors and so on are not giving you the same boost that we were getting 20 to 10 years ago. But I think at the same time, we're seeing much more specialized computational devices, like machine learning accelerators, TPUs, and very ML-focused GPUs, more recently, are making it so that we can actually get really high performance and good efficiency out of the more modern kinds of computations we want to run that are different than a twisty pile of C++ code trying to run Microsoft Office or something.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

53 / evaluation

Well, I would say that the pivot to hardware oriented around that was an important transition, because before that, we had CPUs and GPUs that were not especially well-suited for deep learning.

“Well, I would say that the pivot to hardware oriented around that was an important transition, because before that, we had CPUs and GPUs that were not especially well-suited for deep learning.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

54 / evaluation

I think one of the things we were a little, our view of things from a search perspective was these models hallucinate a lot, they don't get things right a lot of the time- or some of the time- and that means that they aren't as useful as they could be and so we’d like to make that better.

“I think one of the things we were a little, our view of things from a search perspective was these models hallucinate a lot, they don't get things right a lot of the time- or some of the time- and that means that they aren't as useful as they could be and so we’d like to make that better.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

55 / prediction

I thought, naive me, that 32 processors would be able to train really awesome neural nets. But it turned out we needed about a million times more compute before they really started to work for real problems, but then starting in the late 2008, 2009, 2010 timeframe, we started to have enough compute, thanks to Moore's law, to actually make neural nets work for real things.

“I thought, naive me, that 32 processors would be able to train really awesome neural nets. But it turned out we needed about a million times more compute before they really started to work for real problems, but then starting in the late 2008, 2009, 2010 timeframe, we started to have enough compute, thanks to Moore's law, to actually make neural nets work for real things.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

56 / prediction

Compute is the rough, highest-level view of these capable models because if one of the techniques for improving their quality is scaling up the amount of inference compute you use, then all of a sudden what's currently like one request to generate some tokens now becomes 50 or 100 or 1000 times as computationally intensive, even though it's producing the same amount of output.

“Compute is the rough, highest-level view of these capable models because if one of the techniques for improving their quality is scaling up the amount of inference compute you use, then all of a sudden what's currently like one request to generate some tokens now becomes 50 or 100 or 1000 times as computationally intensive, even though it's producing the same amount of output.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

57 / prediction

Some of those benefits may be improved quality, some may be less concretely measurable, like this ability to have lots of parallel development of different modules. But that's still a pretty exciting improvement because I think that would enable us to make faster progress on improving the model's capabilities for lots of different distinct areas.

“Some of those benefits may be improved quality, some may be less concretely measurable, like this ability to have lots of parallel development of different modules. But that's still a pretty exciting improvement because I think that would enable us to make faster progress on improving the model's capabilities for lots of different distinct areas.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast
Search evidence