app / likes
ChatGPT app
“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”
Public evidence record
Published podcast speaker
Books, apps, and tools
app / likes
“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”
app / likes
“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”
book / recommends
“It’s a great book, Season of the Witch; I recommend it. A bunch of my SF friends who do get out recommended it to me.”
Claim ledger
117 transcript-backed records
01 / recommendation
“I think I would say, one of the people I worked with is moving to SF, and I need to get him a copy of Season of the Witch. It’s a history of SF from 1960 to 1985 that goes through the hippie revolution, the culture emerging in the city, the HIV/AIDS crisis, and other things. That is so recent, with so much turmoil and hurt, but also love in SF. No one knows about this. It’s a great book, Season of the Witch; I recommend it. A bunch of my SF friends who do get out recommended it to me. I lived there and I didn’t appreciate this context, and it’s just so recent.”
02 / recommendation
“It’s a great book, Season of the Witch; I recommend it. A bunch of my SF friends who do get out recommended it to me.”
03 / belief
“I think in an era when things are moving very fast and are very chaotic, it’s very rewarding to people.”
04 / belief
“I think the question is if the companies can support the valuations. I’d see the AI companies being looked at in some ways like AWS, Azure, and GCP, which are all competing in the same space and all very successful businesses.”
05 / belief
“I think when you look at China, the biggest reason is that they want people around the world to use these models, and I think a lot of people will not.”
06 / belief
“I think a lot of that’s happened at labs this year; there are new hot things, whether it’s coding environments or web navigation, and you just need to bring in new data and change your whole pre-training so that your post-training can work better.”
07 / belief
“I would say less than that on the software side, but I think longer than that on things like research.”
08 / belief
“I think at a lot of frontier labs, when they scale researchers, a lot more goes into data.”
09 / belief
“All the AI interfaces are getting set up to ask humans for input. I think Claude Code we talked about a lot.”
10 / belief
“I think that dream is actually kind of dying. As you talked about with the specialized models where it’s like… and multimodal is often… like, video generation is a totally different thing.”
11 / commitment
“I will regularly have five pro queries going simultaneously, each looking for one specific paper or feedback on an equation.”
12 / belief
“I think that that’s hard but if you want to scope the maximum possible impact with minimum compute, it’s something like that—which is just get very narrow, and it takes learning of where the models are going.”
13 / belief
“I think OpenAI’s definition is somewhat related to that—an AI that can do a certain number of economically valuable tasks—which I don’t really love as a definition, but it could be a grounding point.”
14 / belief
“I think GPUs would still exist… …At the time of AlexNet and at the time of the Transformer.”
15 / belief
“I think eventually they’ll fool you, and it’ll be on platforms that give ways of verifying or building trust.”
16 / belief
“I think you could train models to do this and it would be a wonderful contribution.”
17 / belief
“I think there will be some other big multi-billion dollar acquisitions, like Perplexity.”
18 / belief
“I think that if you’re doing a PhD, you could also be like, “It’s too risky to work in language models.”
19 / belief
“If scaling laws are fundamental in deep learning, I think the Bitter Lesson will always apply, which is compute will become more abundant.”
20 / belief
“I think that there’s—we’ll get to continual learning later, but there’s a lot of buzz around certain areas of AI, but no one knows when the next step function will really come.”
21 / belief
“We saw multiple demos in 2025 of, like, Claude can use your computer, or OpenAI had operator, and they all suck. So they’re investing money in this, and I think that’ll be a good example.”
22 / belief
“I think if the organizations allow it, AI could very easily implement features end-to-end and do a fairly good job for things that you want to try.”
23 / belief
“Eventually someone else will still have the idea. So I think that in that way, Jensen is helping manifest this GPU revolution much faster and much more focused than it would be without having a person like him there.”
24 / belief
“I think if you think of it over 100 years, society can be changed more with more compute and intelligence because of autonomy.”
25 / uncertainty
“I guess I don’t know— how bad production codebases are, but I think that within… on the order of a few years, a lot of people are going to be pushed to be more like a designer and product manager, where you have multiple of these agents that can try things for you, and they might take one to two days to implement a feature or attempt to fix a bug.”
26 / belief
“On the Manhattan Project thing, one of my funny things looking at them is I think that a Manhattan Project-like thing for open models would actually be pretty reasonable, because it wouldn’t cost that much.”
27 / belief
“Whenever they make a release, they’re always talking about how their GPUs are hurting. And I think in one of these gpt-oss-120b release sessions, Sam Altman said, “Oh, we’re releasing this because we can use your GPUs.”
28 / belief
“I think that when they’re close to this automated software engineer, what it will be good at is traditional ML systems and front end—the model is excellent at those—but the distributed ML, the models are actually really quite bad at because there’s so little training data on doing large-scale distributed learning and things.”
29 / uncertainty
“A lot of it is financial services, so I don’t know what this is. It’s just hard for me to think about the GDP bump, but I would say that software development becomes valuable in a different way when you no longer have to look at the code anymore.”
30 / belief
“Anthropic, I think, bought thousands of books and scanned them and was cleared legally for that because they bought the books, and that is going through the system.”
31 / belief
“I think process reward models were tried a lot more in the pre-o1 era, and a lot of people had headaches with them.”
32 / belief
“I think we’ll still carry around a physical brick of compute— —because people want some ability to have a private interface.”
33 / uncertainty
“Unless they were great, but the first ads won’t be great because it’s a hard problem that we don’t know how to solve.”
34 / belief
“I think the deal for Groq to NVIDIA is rumored to be better for the employees, but it is still this antitrust-avoiding thing.”
35 / belief
“I think the leap from the AI singularity to scaling up mass manufacturing in the US because we have a massive AI advantage is one that is troubled by a lot of political and other challenging problems.”
36 / belief
“I think of the continual learning thing as a research problem where there could be a breakthrough that makes transformers work way better at this and it’s cheap.”
37 / belief
“Linking to what’s happening in big tech, this AI 2027 report leans into the singularity idea where I think research is messy and social and largely in the data in ways that AI models can’t process.”
38 / commitment
“I think we will. I’m definitely a worrier both about AI and non-AI things, but humans do tend to find a way.”
39 / belief
“I think that all of your decisions when you’re training a model come back to pre-training.”
40 / evaluation
“There are now tons of tech companies in China that are releasing very strong frontier open weight models, to the point where I would say that DeepSeek is kind of losing its crown as the preeminent open model maker in China, and the likes of Z.”
41 / prediction
“I think that open models will struggle to replicate some of the things that I like to do with closed models, where you can reference a mix of public and private information.”
42 / evaluation
“I think the tool use point is the one that’s stopping them from being most general purpose because, with something like Claude Code or ChatGPT with search, the autoregressive chain is interrupted with an external tool, and I don’t know how to do that with the diffusion setup.”
43 / preference
“This is why I like the ChatGPT app, because it gives the AI a home in your computer where you can focus on it, rather than just being another tab in my mess of internet options.”
44 / evaluation
“I think in education, a lot of it needs to be, at this point, what I like— —because language models are so good at the math.”
45 / evaluation
“While I’m talking, I’ll say that the Chinese open language models tend to be much bigger and that gives them this higher peak performance as MoEs, whereas a lot of these things that we like a lot, whether it was Gemma or Nemotron, have tended to be smaller models from the US, which is starting to change.”
46 / commitment
“The idea is, if we are going to get to something that is a true, general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-the-job learning.”
47 / preference
“I will always edge on that side when the progress is very high because you don’t know when that’ll unlock a new use case.”
48 / evaluation
“The dense model holds most of the weights if you count them in a transformer model, so you can get really big gains from those mixture of experts on parameter efficiency at training and inference because you get this efficiency by not activating all of these parameters.”
49 / uncertainty
“Is running in parallel actually search? Because I don’t know if we have the full information on how o1-pro works.”
50 / belief
“I think we need to make a point clear on why the time is now for people that don’t think about this, because essentially, with export controls, you’re making it so China cannot make or get cutting edge chips.”
51 / belief
“I think if you have the demand and the money is on the line, the American companies figure it out. It’s going to take handholding with the government, but I think that the culture helps TSMC break through and it’s easier for them.”
52 / uncertainty
“The thing that amplifies the relevance of culture with language models is that we are used to this mode of interacting with people in back and forth conversation. And we now have very powerful computer system that slots into a social context that we’re used to, which makes people very… We don’t know the extent that which people can be impacted by that.”
53 / belief
“There’s a lot of really specific things you can do, but all of this is about fine-tuning to human preferences. And the final stage is much newer and will link to what is done in R1 and these reasoning models is I think OpenAI’s name for this, they had this new API in the fall, which they called the reinforcement fine-tuning API.”
54 / belief
“I have XYZ constraints,” and actually trusting it. I think there’s an HCI problem coming back for information.”
55 / belief
“I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.”
56 / commitment
“I think the US has made it clear to Chinese leaders that we intend to control this technology at whatever cost to global economic integration.”
57 / uncertainty
“Many companies have done seven nanometer chips. And the question is we don’t know how much Huawei was subsidizing production of that chip.”
58 / belief
“I think there’s been some people that are higher level economics understanding say that as you go from 1 billion of smuggling to 10 billion, it’s like you’re hiding certain levels of economic activity and that’s the most reasonable thing to me is that there’s going to be some level where it’s so obvious that it’s easier to find this economic activity.”
59 / belief
“There are other international things that are worrying, but there’s just fundamental human goodness and trying to amplify that. I think we’re on a tenuous time.”
60 / uncertainty
“” The actual thing that happened is much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade, I don’t know anything about financial markets, but it builds up over the weekend, the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming, but it comes on the weekend and then everybody wants to sell, and then that is a social contagion.”
61 / uncertainty
“I would say the simplest one is that our language models to date have been designed to give the right answer the highest percentage of the time in one response. And we are now opening the door to different ways of running inference on our models in which we need to reevaluate many parts of the training process, which normally opens the door to more progress, but we don’t know if OpenAI changed a lot or if just sampling more and multiple choice is what they’re doing or if it’s something more complex, but they changed the training and they know that the inference mode is going to be different.”
62 / belief
“Even things that if you’re an expert, things that are close to the fringe of knowledge, they will still be fairly good at, I think.”
63 / belief
“I think a lot of the AI industry is going through this challenge of communications right now where OpenAI makes fun of their own naming schemes.”
64 / observation
“Character AI very likely could be optimizing this where it’s the way that this data is collected is naive, whereas you’re presented a few options and you choose them. But that’s not the only way that these models are going to be trained.”
65 / belief
“I think their products have long since been banned in China, and I respect saying it directly.”
66 / belief
“I think language models are a form of AGI and all of this super powerful stuff is a next step that’s great if we get these tools.”
67 / belief
“I think then it’ll reassess of what is the biggest problem facing AI and tack on a different angle to the wild ride that we’re on.”
68 / belief
“I think the clearest example we have, because Meta is also open, they talk about order of 60k to 100k H100 equivalent GPUs in their training clusters.”
69 / belief
“The big picture is that I don’t think it’s going to be a cliff. I think a really good example of how growth changes is when Meta added stories.”
70 / belief
“I think humans will definitely be around in a 1000 years, I think. There’s ways that very bad things could happen.”
71 / belief
“I would say that the long tail of use is going to go inside of AI, which is if you scrape trillions of tokens of data, you’re not looking and saying, “This one New York Times article is so important to me.”
72 / uncertainty
“I think for reasoning with this RL and verifiable domains, we’re early, but we don’t know where the point is where you just start training on enough domains and poof, more domains just start working.”
73 / belief
“We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.”
74 / belief
“I think we should comment the why Chinese economy would be hurt by that is that they’re export heavy, I think.”
75 / belief
“For the past few years, the highest cost human data has been in these preferences, which is comparing, I would say, highest cost and highest total usage, so a lot of money has gone to these pairwise comparisons where you have two model outputs and a human is comparing between the two of them.”
76 / belief
“I think export controls are decapping the amount of compute or the density of compute that China can have.”
77 / belief
“I think in terms of internet posts and things that people have been measuring, it hasn’t been a exponential increase or something extremely measurable and things you’re talking about with voice calls and stuff like that, it could be in modalities that are harder to measure.”
78 / belief
“Very long-term motivated in how the ecosystem of AI should work. And I think from a Chinese perspective, he wants a Chinese company to build this vision.”
79 / evaluation
“Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.”
80 / evaluation
“If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.”
81 / prediction
“I think we should summarize what The Bitter Lesson actually is about, is that The Bitter Lesson essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that which are scalable in learning and search, is what it calls out.”
82 / prediction
“I think mostly if you’re going to say that I’m feeling the AGI is that I expect continued, rapid, surprising progress over the next few years. So, something like R1 is less surprising to me from DeepSeek because I expect there to be new paradigms versus … … surprising to me from DeepSeek because I expect there to be new paradigms where substantial progress can be made.”
83 / evaluation
“There’s not full agreement in the community, but for us that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.”
84 / evaluation
“Because the model weights for DeepSeek-R1 are openly available and the license is very friendly, the MIT license commercially available, all of these midsize companies and big companies are trying to be first to serve R1 to their users.”
85 / evaluation
“We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software.”
86 / evaluation
“As reinforcement learning is so much less compute, like it is a richer signal in terms of its impact. Because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner.”
87 / belief
“Like I know some that are in the like the RLHF as a service space will become busy. I think for good reason, just because like.”
88 / belief
“I think if you zoom into any of the details to look at like the agreement number, so how if you look at a test set, you'll have a chosen and rejected and you can take the reward model you're training, pass in those completions and you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward.”
89 / belief
“Oh, I don't even know what inst means, but just saying like they use their adjective that they like. I think Entropic also like steerable is another one.”
90 / belief
“I think big labs are indexed on their own base models so they don't know like what's swapping between CloudBase or GPT-4 base how that would change any notion of preference or what you do with RLHF.”
91 / belief
“I think the things that people see now is like the small models don't really handle nuance as well and they could be more repetitive even if they have really good instruction tuning.”
92 / belief
“I think their papers are sometimes pretty funny because they're not capabilities papers.”
93 / belief
“I think with the language model, it's very hard to define what an environment is.”
94 / belief
“I think the reason why it's not really talked about is just because the RLHF techniques that people use were built in labs like OpenAI and DeepMind where there are some of these people.”
95 / uncertainty
“I try to do it every six or 12 months is my estimated cadence, just to refine the ways that I say things. And people will see that we don't know that much more, but we have a bit of better way of saying what we don't know.”
96 / uncertainty
“People release checkpoints, but that's how we should be thinking about it because the optimizer is so strong and it's like we don't know what's happening in this kind of intermediate land.”
97 / belief
“I think if the people are kind of locked into using synthetic data, people also think that synthetic data is like GPT-4 is more accurate than humans at labeling preferences.”
98 / belief
“I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later.”
99 / belief
“I think I saw people criticizing it for like just being like safety washing from the fact that they're like talking about GPT-2 still, which is such a kind of like odd model to focus on.”
100 / evaluation
“It's really tricky to actually do that. I think that people just keep using GPT-4 because it's really cheap.”
101 / evaluation
“I think InstructGPT does something where they like try to get the RL model to match the instruction tuning dataset because they were really happy with that dataset to constrain the distribution.”
102 / evaluation
“I think the reason this really is done on a deep level is that you're not actually trying to model any contestable preference in this.”
103 / prediction
“I think in the long run, it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about.”
104 / evaluation
“Is this text bad? That's not that surprising, I think, because you could use like a hundred times smaller language model and do much better at filtering than RLHF.”
105 / evaluation
“However, reinforcement learning proved highly effective, particularly given its cost and time effectiveness. So you don't really know exactly what the costs and time that Meta is looking at, because they have a huge team and a pretty good amount of money here to release these Llama models.”
106 / commitment
“There's an early one on RLHF, which is, this stuff is all just like when I figure it out in my brain. So I wrote an article that's like how RLHF actually works, which is just the intuitions that I had put together in the summer about RLHF, and that was pretty well.”
107 / evaluation
“Like, I don't know exactly if they're doing this now, but you can kind of see why doing RLHF at scale and prioritizing a lot of different endpoints would be hard because these are all things I'd be interested in if I was scaling up a big team to do RLHF and like what is going into the preference data.”
108 / evaluation
“There's essentially the thing in my mind that I can't get past is the difference between the control you get in training a reward model and then training a policy because essentially everything you want your reward model to do might not be everything that you train the policy to do in the RLHF step where you have like the two different prompt distributions.”
109 / recommendation
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche, because you're not going to make your model chat GPT like better than OpenAI or anything like that.”
110 / belief
“Like hugging face. I think every. Library, like all these people at Hugging and Face, were working super hard this weekend to make day zero support for Llama2.”
111 / uncertainty
“Near, especially after three organizational restructures of researchers hopping, playing hopscotch from one org to another, and being in between, in between jobs. I don't know.”
112 / belief
“Generally we're trying to operate in the scale a little bit smaller than what Meta is doing cuz we obviously don't have that kind of resources at a startup. So I do a lot of technical research and also try to actually engage and communicate that with the community and specifically, Llama, I think I was most interested on kind of the research side.”
113 / belief
“I think may like, it seems like now that both surge and scale are claiming some part in it, which I find hilarious.”
114 / belief
“I, I think the commercial use of this is gonna be off the charts very soon, like at hugging face.”
115 / recommendation
“We like did a, like a nice experimentation of that hugging face and it's, it's out there, it's ready for someone to invest more time in it and do it.”
116 / prediction
“I like, that's what everyone in my circles is saying is the trend and given machine learning in the last few years, I think trends tend to be stickier than most people expect them to be.”
117 / evaluation
“I think technical terms are like deduplication, so you don't wanna pass the model, the same text, even if it came from different websites and there's tons more that goes into this.”