High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Nathan Lambert

Published podcast speaker

Claims
117
Episodes
4
Shows
2
Named items
3

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

app / likes

ChatGPT app

“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”

Lex Fridman Podcast · 1 Feb 2026

Evidence receipt · Source ↗

app / likes

ChatGPT

“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”

Lex Fridman Podcast · 1 Feb 2026

Evidence receipt · Source ↗

book / recommends

Season of the Witch

“It’s a great book, Season of the Witch; I recommend it. A bunch of my SF friends who do get out recommended it to me.”

Lex Fridman Podcast · 1 Feb 2026

Evidence receipt · Source ↗

Claim ledger

What Nathan said.

117 transcript-backed records

01 / recommendation

Recommends Season of the Witch.

“I think I would say, one of the people I worked with is moving to SF, and I need to get him a copy of Season of the Witch. It’s a history of SF from 1960 to 1985 that goes through the hippie revolution, the culture emerging in the city, the HIV/AIDS crisis, and other things. That is so recent, with so much turmoil and hurt, but also love in SF. No one knows about this. It’s a great book, Season of the Witch; I recommend it. A bunch of my SF friends who do get out recommended it to me. I lived there and I didn’t appreciate this context, and it’s just so recent.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

04 / belief

I think the question is if the companies can support the valuations. I’d see the AI companies being looked at in some ways like AWS, Azure, and GCP, which are all competing in the same space and all very successful businesses.

“I think the question is if the companies can support the valuations. I’d see the AI companies being looked at in some ways like AWS, Azure, and GCP, which are all competing in the same space and all very successful businesses.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

05 / belief

I think when you look at China, the biggest reason is that they want people around the world to use these models, and I think a lot of people will not.

“I think when you look at China, the biggest reason is that they want people around the world to use these models, and I think a lot of people will not.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

06 / belief

I think a lot of that’s happened at labs this year; there are new hot things, whether it’s coding environments or web navigation, and you just need to bring in new data and change your whole pre-training so that your post-training can work better.

“I think a lot of that’s happened at labs this year; there are new hot things, whether it’s coding environments or web navigation, and you just need to bring in new data and change your whole pre-training so that your post-training can work better.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

10 / belief

I think that dream is actually kind of dying. As you talked about with the specialized models where it’s like… and multimodal is often… like, video generation is a totally different thing.

“I think that dream is actually kind of dying. As you talked about with the specialized models where it’s like… and multimodal is often… like, video generation is a totally different thing.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

12 / belief

I think that that’s hard but if you want to scope the maximum possible impact with minimum compute, it’s something like that—which is just get very narrow, and it takes learning of where the models are going.

“I think that that’s hard but if you want to scope the maximum possible impact with minimum compute, it’s something like that—which is just get very narrow, and it takes learning of where the models are going.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

13 / belief

I think OpenAI’s definition is somewhat related to that—an AI that can do a certain number of economically valuable tasks—which I don’t really love as a definition, but it could be a grounding point.

“I think OpenAI’s definition is somewhat related to that—an AI that can do a certain number of economically valuable tasks—which I don’t really love as a definition, but it could be a grounding point.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

20 / belief

I think that there’s—we’ll get to continual learning later, but there’s a lot of buzz around certain areas of AI, but no one knows when the next step function will really come.

“I think that there’s—we’ll get to continual learning later, but there’s a lot of buzz around certain areas of AI, but no one knows when the next step function will really come.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

21 / belief

We saw multiple demos in 2025 of, like, Claude can use your computer, or OpenAI had operator, and they all suck. So they’re investing money in this, and I think that’ll be a good example.

“We saw multiple demos in 2025 of, like, Claude can use your computer, or OpenAI had operator, and they all suck. So they’re investing money in this, and I think that’ll be a good example.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

23 / belief

Eventually someone else will still have the idea. So I think that in that way, Jensen is helping manifest this GPU revolution much faster and much more focused than it would be without having a person like him there.

“Eventually someone else will still have the idea. So I think that in that way, Jensen is helping manifest this GPU revolution much faster and much more focused than it would be without having a person like him there.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

25 / uncertainty

I guess I don’t know— how bad production codebases are, but I think that within… on the order of a few years, a lot of people are going to be pushed to be more like a designer and product manager, where you have multiple of these agents that can try things for you, and they might take one to two days to implement a feature or attempt to fix a bug.

“I guess I don’t know— how bad production codebases are, but I think that within… on the order of a few years, a lot of people are going to be pushed to be more like a designer and product manager, where you have multiple of these agents that can try things for you, and they might take one to two days to implement a feature or attempt to fix a bug.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

26 / belief

On the Manhattan Project thing, one of my funny things looking at them is I think that a Manhattan Project-like thing for open models would actually be pretty reasonable, because it wouldn’t cost that much.

“On the Manhattan Project thing, one of my funny things looking at them is I think that a Manhattan Project-like thing for open models would actually be pretty reasonable, because it wouldn’t cost that much.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

27 / belief

Whenever they make a release, they’re always talking about how their GPUs are hurting. And I think in one of these gpt-oss-120b release sessions, Sam Altman said, “Oh, we’re releasing this because we can use your GPUs.

“Whenever they make a release, they’re always talking about how their GPUs are hurting. And I think in one of these gpt-oss-120b release sessions, Sam Altman said, “Oh, we’re releasing this because we can use your GPUs.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

28 / belief

I think that when they’re close to this automated software engineer, what it will be good at is traditional ML systems and front end—the model is excellent at those—but the distributed ML, the models are actually really quite bad at because there’s so little training data on doing large-scale distributed learning and things.

“I think that when they’re close to this automated software engineer, what it will be good at is traditional ML systems and front end—the model is excellent at those—but the distributed ML, the models are actually really quite bad at because there’s so little training data on doing large-scale distributed learning and things.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

29 / uncertainty

A lot of it is financial services, so I don’t know what this is. It’s just hard for me to think about the GDP bump, but I would say that software development becomes valuable in a different way when you no longer have to look at the code anymore.

“A lot of it is financial services, so I don’t know what this is. It’s just hard for me to think about the GDP bump, but I would say that software development becomes valuable in a different way when you no longer have to look at the code anymore.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

30 / belief

Anthropic, I think, bought thousands of books and scanned them and was cleared legally for that because they bought the books, and that is going through the system.

“Anthropic, I think, bought thousands of books and scanned them and was cleared legally for that because they bought the books, and that is going through the system.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

35 / belief

I think the leap from the AI singularity to scaling up mass manufacturing in the US because we have a massive AI advantage is one that is troubled by a lot of political and other challenging problems.

“I think the leap from the AI singularity to scaling up mass manufacturing in the US because we have a massive AI advantage is one that is troubled by a lot of political and other challenging problems.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

36 / belief

I think of the continual learning thing as a research problem where there could be a breakthrough that makes transformers work way better at this and it’s cheap.

“I think of the continual learning thing as a research problem where there could be a breakthrough that makes transformers work way better at this and it’s cheap.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

37 / belief

Linking to what’s happening in big tech, this AI 2027 report leans into the singularity idea where I think research is messy and social and largely in the data in ways that AI models can’t process.

“Linking to what’s happening in big tech, this AI 2027 report leans into the singularity idea where I think research is messy and social and largely in the data in ways that AI models can’t process.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

40 / evaluation

There are now tons of tech companies in China that are releasing very strong frontier open weight models, to the point where I would say that DeepSeek is kind of losing its crown as the preeminent open model maker in China, and the likes of Z.

“There are now tons of tech companies in China that are releasing very strong frontier open weight models, to the point where I would say that DeepSeek is kind of losing its crown as the preeminent open model maker in China, and the likes of Z.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

41 / prediction

I think that open models will struggle to replicate some of the things that I like to do with closed models, where you can reference a mix of public and private information.

“I think that open models will struggle to replicate some of the things that I like to do with closed models, where you can reference a mix of public and private information.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

42 / evaluation

I think the tool use point is the one that’s stopping them from being most general purpose because, with something like Claude Code or ChatGPT with search, the autoregressive chain is interrupted with an external tool, and I don’t know how to do that with the diffusion setup.

“I think the tool use point is the one that’s stopping them from being most general purpose because, with something like Claude Code or ChatGPT with search, the autoregressive chain is interrupted with an external tool, and I don’t know how to do that with the diffusion setup.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

43 / preference

This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.

“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

45 / evaluation

While I’m talking, I’ll say that the Chinese open language models tend to be much bigger and that gives them this higher peak performance as MoEs, whereas a lot of these things that we like a lot, whether it was Gemma or Nemotron, have tended to be smaller models from the US, which is starting to change.

“While I’m talking, I’ll say that the Chinese open language models tend to be much bigger and that gives them this higher peak performance as MoEs, whereas a lot of these things that we like a lot, whether it was Gemma or Nemotron, have tended to be smaller models from the US, which is starting to change.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

46 / commitment

The idea is, if we are going to get to something that is a true, general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-the-job learning.

“The idea is, if we are going to get to something that is a true, general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-the-job learning.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

48 / evaluation

The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.

“The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

50 / belief

I think we need to make a point clear on why the time is now for people that don’t think about this, because essentially, with export controls, you’re making it so China cannot make or get cutting edge chips.

“I think we need to make a point clear on why the time is now for people that don’t think about this, because essentially, with export controls, you’re making it so China cannot make or get cutting edge chips.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

51 / belief

I think if you have the demand and the money is on the line, the American companies figure it out. It’s going to take handholding with the government, but I think that the culture helps TSMC break through and it’s easier for them.

“I think if you have the demand and the money is on the line, the American companies figure it out. It’s going to take handholding with the government, but I think that the culture helps TSMC break through and it’s easier for them.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

52 / uncertainty

The thing that amplifies the relevance of culture with language models is that we are used to this mode of interacting with people in back and forth conversation. And we now have very powerful computer system that slots into a social context that we’re used to, which makes people very… We don’t know the extent that which people can be impacted by that.

“The thing that amplifies the relevance of culture with language models is that we are used to this mode of interacting with people in back and forth conversation. And we now have very powerful computer system that slots into a social context that we’re used to, which makes people very… We don’t know the extent that which people can be impacted by that.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

53 / belief

There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.

“There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

55 / belief

I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.

“I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

58 / belief

I think there’s been some people that are higher level economics understanding say that as you go from 1 billion of smuggling to 10 billion, it’s like you’re hiding certain levels of economic activity and that’s the most reasonable thing to me is that there’s going to be some level where it’s so obvious that it’s easier to find this economic activity.

“I think there’s been some people that are higher level economics understanding say that as you go from 1 billion of smuggling to 10 billion, it’s like you’re hiding certain levels of economic activity and that’s the most reasonable thing to me is that there’s going to be some level where it’s so obvious that it’s easier to find this economic activity.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

59 / belief

There are other international things that are worrying, but there’s just fundamental human goodness and trying to amplify that. I think we’re on a tenuous time.

“There are other international things that are worrying, but there’s just fundamental human goodness and trying to amplify that. I think we’re on a tenuous time.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

60 / uncertainty

” The actual thing that happened is much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade, I don’t know anything about financial markets, but it builds up over the weekend, the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming, but it comes on the weekend and then everybody wants to sell, and then that is a social contagion.

“” The actual thing that happened is much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade, I don’t know anything about financial markets, but it builds up over the weekend, the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming, but it comes on the weekend and then everybody wants to sell, and then that is a social contagion.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

61 / uncertainty

I would say the simplest one is that our language models to date have been designed to give the right answer the highest percentage of the time in one response. And we are now opening the door to different ways of running inference on our models in which we need to reevaluate many parts of the training process, which normally opens the door to more progress, but we don’t know if OpenAI changed a lot or if just sampling more and multiple choice is what they’re doing or if it’s something more complex, but they changed the training and they know that the inference mode is going to be different.

“I would say the simplest one is that our language models to date have been designed to give the right answer the highest percentage of the time in one response. And we are now opening the door to different ways of running inference on our models in which we need to reevaluate many parts of the training process, which normally opens the door to more progress, but we don’t know if OpenAI changed a lot or if just sampling more and multiple choice is what they’re doing or if it’s something more complex, but they changed the training and they know that the inference mode is going to be different.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

64 / observation

Character AI very likely could be optimizing this where it’s the way that this data is collected is naive, whereas you’re presented a few options and you choose them. But that’s not the only way that these models are going to be trained.

“Character AI very likely could be optimizing this where it’s the way that this data is collected is naive, whereas you’re presented a few options and you choose them. But that’s not the only way that these models are going to be trained.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

68 / belief

I think the clearest example we have, because Meta is also open, they talk about order of 60k to 100k H100 equivalent GPUs in their training clusters.

“I think the clearest example we have, because Meta is also open, they talk about order of 60k to 100k H100 equivalent GPUs in their training clusters.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

71 / belief

I would say that the long tail of use is going to go inside of AI, which is if you scrape trillions of tokens of data, you’re not looking and saying, “This one New York Times article is so important to me.

“I would say that the long tail of use is going to go inside of AI, which is if you scrape trillions of tokens of data, you’re not looking and saying, “This one New York Times article is so important to me.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

72 / uncertainty

I think for reasoning with this RL and verifiable domains, we’re early, but we don’t know where the point is where you just start training on enough domains and poof, more domains just start working.

“I think for reasoning with this RL and verifiable domains, we’re early, but we don’t know where the point is where you just start training on enough domains and poof, more domains just start working.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

73 / belief

We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.

“We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

75 / belief

For the past few years, the highest cost human data has been in these preferences, which is comparing, I would say, highest cost and highest total usage, so a lot of money has gone to these pairwise comparisons where you have two model outputs and a human is comparing between the two of them.

“For the past few years, the highest cost human data has been in these preferences, which is comparing, I would say, highest cost and highest total usage, so a lot of money has gone to these pairwise comparisons where you have two model outputs and a human is comparing between the two of them.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

77 / belief

I think in terms of internet posts and things that people have been measuring, it hasn’t been a exponential increase or something extremely measurable and things you’re talking about with voice calls and stuff like that, it could be in modalities that are harder to measure.

“I think in terms of internet posts and things that people have been measuring, it hasn’t been a exponential increase or something extremely measurable and things you’re talking about with voice calls and stuff like that, it could be in modalities that are harder to measure.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

78 / belief

Very long-term motivated in how the ecosystem of AI should work. And I think from a Chinese perspective, he wants a Chinese company to build this vision.

“Very long-term motivated in how the ecosystem of AI should work. And I think from a Chinese perspective, he wants a Chinese company to build this vision.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

79 / evaluation

Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.

“Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

80 / evaluation

If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.

“If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

81 / prediction

I think we should summarize what The Bitter Lesson actually is about, is that The Bitter Lesson essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that which are scalable in learning and search, is what it calls out.

“I think we should summarize what The Bitter Lesson actually is about, is that The Bitter Lesson essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that which are scalable in learning and search, is what it calls out.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

82 / prediction

I think mostly if you’re going to say that I’m feeling the AGI is that I expect continued, rapid, surprising progress over the next few years. So, something like R1 is less surprising to me from DeepSeek because I expect there to be new paradigms versus … … surprising to me from DeepSeek because I expect there to be new paradigms where substantial progress can be made.

“I think mostly if you’re going to say that I’m feeling the AGI is that I expect continued, rapid, surprising progress over the next few years. So, something like R1 is less surprising to me from DeepSeek because I expect there to be new paradigms versus … … surprising to me from DeepSeek because I expect there to be new paradigms where substantial progress can be made.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

83 / evaluation

There’s not full agreement in the community, but for us that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.

“There’s not full agreement in the community, but for us that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

84 / evaluation

Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.

“Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

85 / evaluation

We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.

“We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.”
Speaker
Nathan Lambert
Publisher
Lex Fridman Podcast

86 / evaluation

As reinforcement learning is so much less compute, like it is a richer signal in terms of its impact. Because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner.

“As reinforcement learning is so much less compute, like it is a richer signal in terms of its impact. Because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner.”
Speaker
Nathan Lambert
Publisher
Latent Space

88 / belief

I think if you zoom into any of the details to look at like the agreement number, so how if you look at a test set, you'll have a chosen and rejected and you can take the reward model you're training, pass in those completions and you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward.

“I think if you zoom into any of the details to look at like the agreement number, so how if you look at a test set, you'll have a chosen and rejected and you can take the reward model you're training, pass in those completions and you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward.”
Speaker
Nathan Lambert
Publisher
Latent Space

90 / belief

I think big labs are indexed on their own base models so they don't know like what's swapping between CloudBase or GPT-4 base how that would change any notion of preference or what you do with RLHF.

“I think big labs are indexed on their own base models so they don't know like what's swapping between CloudBase or GPT-4 base how that would change any notion of preference or what you do with RLHF.”
Speaker
Nathan Lambert
Publisher
Latent Space

91 / belief

I think the things that people see now is like the small models don't really handle nuance as well and they could be more repetitive even if they have really good instruction tuning.

“I think the things that people see now is like the small models don't really handle nuance as well and they could be more repetitive even if they have really good instruction tuning.”
Speaker
Nathan Lambert
Publisher
Latent Space

94 / belief

I think the reason why it's not really talked about is just because the RLHF techniques that people use were built in labs like OpenAI and DeepMind where there are some of these people.

“I think the reason why it's not really talked about is just because the RLHF techniques that people use were built in labs like OpenAI and DeepMind where there are some of these people.”
Speaker
Nathan Lambert
Publisher
Latent Space

95 / uncertainty

I try to do it every six or 12 months is my estimated cadence, just to refine the ways that I say things. And people will see that we don't know that much more, but we have a bit of better way of saying what we don't know.

“I try to do it every six or 12 months is my estimated cadence, just to refine the ways that I say things. And people will see that we don't know that much more, but we have a bit of better way of saying what we don't know.”
Speaker
Nathan Lambert
Publisher
Latent Space

96 / uncertainty

People release checkpoints, but that's how we should be thinking about it because the optimizer is so strong and it's like we don't know what's happening in this kind of intermediate land.

“People release checkpoints, but that's how we should be thinking about it because the optimizer is so strong and it's like we don't know what's happening in this kind of intermediate land.”
Speaker
Nathan Lambert
Publisher
Latent Space

97 / belief

I think if the people are kind of locked into using synthetic data, people also think that synthetic data is like GPT-4 is more accurate than humans at labeling preferences.

“I think if the people are kind of locked into using synthetic data, people also think that synthetic data is like GPT-4 is more accurate than humans at labeling preferences.”
Speaker
Nathan Lambert
Publisher
Latent Space

98 / belief

I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later.

“I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later.”
Speaker
Nathan Lambert
Publisher
Latent Space

99 / belief

I think I saw people criticizing it for like just being like safety washing from the fact that they're like talking about GPT-2 still, which is such a kind of like odd model to focus on.

“I think I saw people criticizing it for like just being like safety washing from the fact that they're like talking about GPT-2 still, which is such a kind of like odd model to focus on.”
Speaker
Nathan Lambert
Publisher
Latent Space

101 / evaluation

I think InstructGPT does something where they like try to get the RL model to match the instruction tuning dataset because they were really happy with that dataset to constrain the distribution.

“I think InstructGPT does something where they like try to get the RL model to match the instruction tuning dataset because they were really happy with that dataset to constrain the distribution.”
Speaker
Nathan Lambert
Publisher
Latent Space

103 / prediction

I think in the long run, it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about.

“I think in the long run, it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about.”
Speaker
Nathan Lambert
Publisher
Latent Space

104 / evaluation

Is this text bad? That's not that surprising, I think, because you could use like a hundred times smaller language model and do much better at filtering than RLHF.

“Is this text bad? That's not that surprising, I think, because you could use like a hundred times smaller language model and do much better at filtering than RLHF.”
Speaker
Nathan Lambert
Publisher
Latent Space

105 / evaluation

However, reinforcement learning proved highly effective, particularly given its cost and time effectiveness. So you don't really know exactly what the costs and time that Meta is looking at, because they have a huge team and a pretty good amount of money here to release these Llama models.

“However, reinforcement learning proved highly effective, particularly given its cost and time effectiveness. So you don't really know exactly what the costs and time that Meta is looking at, because they have a huge team and a pretty good amount of money here to release these Llama models.”
Speaker
Nathan Lambert
Publisher
Latent Space

106 / commitment

There's an early one on RLHF, which is, this stuff is all just like when I figure it out in my brain. So I wrote an article that's like how RLHF actually works, which is just the intuitions that I had put together in the summer about RLHF, and that was pretty well.

“There's an early one on RLHF, which is, this stuff is all just like when I figure it out in my brain. So I wrote an article that's like how RLHF actually works, which is just the intuitions that I had put together in the summer about RLHF, and that was pretty well.”
Speaker
Nathan Lambert
Publisher
Latent Space

107 / evaluation

Like, I don't know exactly if they're doing this now, but you can kind of see why doing RLHF at scale and prioritizing a lot of different endpoints would be hard because these are all things I'd be interested in if I was scaling up a big team to do RLHF and like what is going into the preference data.

“Like, I don't know exactly if they're doing this now, but you can kind of see why doing RLHF at scale and prioritizing a lot of different endpoints would be hard because these are all things I'd be interested in if I was scaling up a big team to do RLHF and like what is going into the preference data.”
Speaker
Nathan Lambert
Publisher
Latent Space

108 / evaluation

There's essentially the thing in my mind that I can't get past is the difference between the control you get in training a reward model and then training a policy because essentially everything you want your reward model to do might not be everything that you train the policy to do in the RLHF step where you have like the two different prompt distributions.

“There's essentially the thing in my mind that I can't get past is the difference between the control you get in training a reward model and then training a policy because essentially everything you want your reward model to do might not be everything that you train the policy to do in the RLHF step where you have like the two different prompt distributions.”
Speaker
Nathan Lambert
Publisher
Latent Space

109 / recommendation

I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche, because you're not going to make your model chat GPT like better than OpenAI or anything like that.

“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche, because you're not going to make your model chat GPT like better than OpenAI or anything like that.”
Speaker
Nathan Lambert
Publisher
Latent Space

110 / belief

Like hugging face. I think every. Library, like all these people at Hugging and Face, were working super hard this weekend to make day zero support for Llama2.

“Like hugging face. I think every. Library, like all these people at Hugging and Face, were working super hard this weekend to make day zero support for Llama2.”
Speaker
Nathan Lambert
Publisher
Latent Space

111 / uncertainty

Near, especially after three organizational restructures of researchers hopping, playing hopscotch from one org to another, and being in between, in between jobs. I don't know.

“Near, especially after three organizational restructures of researchers hopping, playing hopscotch from one org to another, and being in between, in between jobs. I don't know.”
Speaker
Nathan Lambert
Publisher
Latent Space

112 / belief

Generally we're trying to operate in the scale a little bit smaller than what Meta is doing cuz we obviously don't have that kind of resources at a startup. So I do a lot of technical research and also try to actually engage and communicate that with the community and specifically, Llama, I think I was most interested on kind of the research side.

“Generally we're trying to operate in the scale a little bit smaller than what Meta is doing cuz we obviously don't have that kind of resources at a startup. So I do a lot of technical research and also try to actually engage and communicate that with the community and specifically, Llama, I think I was most interested on kind of the research side.”
Speaker
Nathan Lambert
Publisher
Latent Space

115 / recommendation

We like did a, like a nice experimentation of that hugging face and it's, it's out there, it's ready for someone to invest more time in it and do it.

“We like did a, like a nice experimentation of that hugging face and it's, it's out there, it's ready for someone to invest more time in it and do it.”
Speaker
Nathan Lambert
Publisher
Latent Space

116 / prediction

I like, that's what everyone in my circles is saying is the trend and given machine learning in the last few years, I think trends tend to be stickier than most people expect them to be.

“I like, that's what everyone in my circles is saying is the trend and given machine learning in the last few years, I think trends tend to be stickier than most people expect them to be.”
Speaker
Nathan Lambert
Publisher
Latent Space

117 / evaluation

I think technical terms are like deduplication, so you don't wanna pass the model, the same text, even if it came from different websites and there's tons more that goes into this.

“I think technical terms are like deduplication, so you don't wanna pass the model, the same text, even if it came from different websites and there's tons more that goes into this.”
Speaker
Nathan Lambert
Publisher
Latent Space
Search evidence