other / uses
Gwern's book reviews
“The book reviews is a good suggestion. I actually use, like, Gwern's book reviews as a way to recommend books to people.”
Public evidence record
Published podcast speaker
Books, apps, and tools
other / uses
“The book reviews is a good suggestion. I actually use, like, Gwern's book reviews as a way to recommend books to people.”
Claim ledger
59 transcript-backed records
01 / belief
“In particular I think the intellectual ceiling goes quite—contra what I was saying before, which is we've demonstrated this incredible complexity of math, and programming problems… I do think that the type of task and setting that AlphaZero worked in this two-player perfect information game basically is incredibly friendly to RL algorithms.”
02 / belief
“I think what we're seeing now is closer to: lack of context, lack of ability to do complex, very multi-file changes… sort of the scope of the task, in some respects.”
03 / belief
“I think at the beginning, the hill to climb. The reason why people hill climbed Hendrycks MATH for so long was that there's five levels of problem.”
04 / belief
“I can't really talk about where exactly people sit on that scaffold. I think different people, different tasks are on different points there.”
05 / belief
“I think an interesting question over the next few years is whether that is totally sufficient, whether this raw base intelligence, plus sufficient scaffolding in text, is enough to build context, or whether you need to somehow update the weights for your use case, or some combination thereof.”
06 / belief
“If you made an extremely efficient transform implementation on TPU, or Trainium, or Incuda, then I think there's a pretty high likelihood that you'll get a job offer.”
07 / belief
“I think of this product exponential in some respects where you need to be designing for a few months ahead of the model, to make sure that the product you build is the right one.”
08 / belief
“We're here working on AI research. I think each of the companies is trying to define this for themselves.”
09 / belief
“I think we're already seeing early evidence of this in its ability to generalize reasoning to things.”
10 / belief
“I think that would be somewhat of an update towards, there's something strangely difficult about this computer use in particular.”
11 / belief
“I think their research taste is good in a way that I think Noam's research taste is good.”
12 / belief
“As you use more compute, and as you train on more, and more difficult tasks, your rate of improvement of biology for example is going to be somewhat bound by the time it takes a cell to grow in a way that your rate of improvement on math isn't, for example. So, yes, but I think for many things we'll be able to parallelize widely enough, and get enough iteration loops.”
13 / belief
“The residual stream is like operating RAM, you're doing stuff to it, is the mental model I think one takes away from interpretability work.”
14 / belief
“I think the LLM curves look a bit different, in that there isn't that dead zone at the beginning.”
15 / belief
“I think in a lot of these cases you have to hope for some amount of generator verifier gap.”
16 / prediction
“I expect scientific areas where you are able to put it in a feedback loop to have, eventually, superhuman performance.”
17 / belief
“I think that now that RL's come back, papers building on Andy Jones's “Scaling scaling laws for board games” are interesting.”
18 / evaluation
“Because a lot of the tasks required in winning a Nobel Prize—or at least strongly assisting in helping to win a Nobel Prize—have more layers of verifiability built up.”
19 / observation
“I think, in general, when people are talking about the separate model… For example, most of the robotics companies are doing this bi-level thing, where they have a motor policy that's running at 60 hertz or whatever, and some higher-level visual language model.”
20 / observation
“Actually, I don't know if it's good or bad, but Meta didn't include it in Llama, but Deepseek did include it in their paper, which I think is interesting.”
21 / prediction
“Holding me accountable for my predictions next year, I really do think by the end of this year to this time next year, we will have software engineering agents that can do close to a day's worth of work for a junior engineer, or a couple of hours of quite competent, independent work.”
22 / prediction
“Because the thing that these models get good at by default is software engineering and computer using agents and this kind of stuff.”
23 / belief
“I think you do that quite proactively in terms of you deliberately introduce people that you think will be interesting to each other and this kind of stuff, so… yeah.”
24 / belief
“I think making sure that, like, there is a great debate of ideas on not, not just AI, but on other fields, and everything is incredibly high leverage in value.”
25 / evaluation
“I think my answer at the moment is that the sort of pre-training objective doesn't necessarily- like it imbues with this nice flexible general knowledge about the world, but doesn't necessarily imbue the skill of making novel connections or research.”
26 / prediction
“" I expect that trend line to continue, basically, as you go from this model of, "Well, I'm working with some model that's assisting me on my computer, and it's basically a pairing session," to, "I'm managing a small team," through to, "I'm managing a division or a company".”
27 / recommendation
“The book reviews is a good suggestion. I actually use, like, Gwern's book reviews as a way to recommend books to people.”
28 / evaluation
“That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.”
29 / belief
“I think a lot of people look internally these days for their sources of insight or progress.”
30 / uncertainty
“I don't know if you guys have read that article about Jeff and Sanjay, but they were there pair programming on stuff.”
31 / belief
“I think that's more about nines of reliability and the model actually successfully doing things.”
32 / belief
“I think a wonderful research project to do, if someone is out there listening to this, would be to try and take some of the techniques that Trenton's team has worked on and try and disentangle the neurons in the Mistral paper, Mixtral model, which is open source.”
33 / belief
“When I think of really good data, to me, that raises something which involved a lot of reasoning to create.”
34 / belief
“I think directions like publishing the constitution that you expect your model to abide by–trying to make sure that you RLHF it towards that, and ablate that, and have the ability for everyone to offer feedback and contribution to that–is really important.”
35 / belief
“You could imagine biology papers going here, math papers going here, and all of a sudden your breakdown is ruined. But that vision transformer one, where the class separation is really clear and obvious, gives I think some evidence towards the specialization hypothesis.”
36 / belief
“I think Sasha Rush has a great tweet where he basically plots the curve of the cost of attention respective to the cost of really large models and attention actually trails off.”
37 / belief
“Maybe not the user, but we should be able to understand and interpret what these values are doing and the information that is transmitting. I think that's a really important goal for the future.”
38 / belief
“I think the Gemini program would probably be maybe five times faster with 10 times more compute or something like that.”
39 / belief
“I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale, and interpreting and understanding what goes wrong.”
40 / belief
“I think that is true to many extents. I'm sure you probably benefited a lot from the key researchers mentoring you deeply.”
41 / commitment
“We just look at all models on a similar data set. We will learn the same features in the same order-ish.”
42 / belief
“I think the original chain-of-thought paper had that as almost an immersion property of the model.”
43 / belief
“At the moment, I think most of the labs are somewhat compute bound in that there are always more experiments you could run and more pieces of information that you could gain in the same way that scientific research on biology is somewhat experimentally throughput-bound.”
44 / belief
“I think that's really worth exploring. For example, one of the evals that we did in the paper had it learn a language in context better than a human expert could, over the course of a couple of months.”
45 / belief
“I think there are a number of people who are really, really critical. If you took them out then the performance of the program would be dramatically impacted.”
46 / belief
“The ruthless prioritization is something which I think separates a lot of quality research from research that doesn't necessarily succeed as much.”
47 / belief
“” I don't think that's quite true because I think in AI research most people actually care quite deeply.”
48 / belief
“I think John Carmack had this nice phrase where it's the first time in history where you can plausibly imagine writing AI with 10,000 lines of code.”
49 / belief
“I think one thing that I didn't fully convey before was that I think a lot of like good research comes from working backwards from the actual problems that you want to solve.”
50 / belief
“” There's also someone who works on Anthropic's performance team now, Simon Boehm, who has written in my mind the reference for optimizing a CUDA map model on a GPU.”
51 / belief
“I think that's arguably the most important quality in almost anything. It's just pursuing it to the end of the earth.”
52 / belief
“There’s one mode of thinking in which fine-tuning is specialized, you've got this latent bundle of capabilities and you're specializing it for this particular use case that you want. I think I'm not sure how true or not that is.”
53 / belief
“I think the traditional story for why distillation is more efficient is during training, normally you're trying to predict this one hot vector that says, “this is the token that you should have predicted.”
54 / belief
“I think there's an important aspect of shots on goal there, so to speak. Where just choosing to go to conferences itself is putting yourself in a position where luck is more likely to happen.”
55 / evaluation
“I think we are less, at the moment, bound by the sheer engineering work of making these things than we are by compute to run and get signal, and taste in terms of what the actual right thing to do is.”
56 / evaluation
“Until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved.”
57 / preference
“I think many people make the decision that the thing that they want to prioritize is a wonderful life with their family.”
58 / evaluation
“I mean problems that haven't been particularly well-solved so far, but perhaps as a result of frustrating structural factors like the ones that you pointed out in that scenario before, where they're like, “we can't do X because this team won’t do Y.”
59 / evaluation
“There's a line of work I quite like, where it looks at in-context learning as basically very similar to gradient descent, but the attention operation can be viewed as gradient descent on the in-context data. That paper had some cool plots where they basically showed “we take n steps of gradient descent and that looks like n layers of in-context learning, and it looks very similar.”