High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

9 Apr 2026 · 46:04 Unsupervised Learning Ep 84: OpenAI’s Chief Scientist on Continual Learning Hype, RL Beyond Code, & Future Alignment Directions

“The the way my thinking about the problem has evolved over the past few years is definitely kind of gone from you know, always this like very nebulous problem that like it's just like very hard to even grapple with or define uh to like, oh, you know, I I think we can actually make progress at it by very concrete technical solutions and technical insights.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
9 Apr 2026 · 46:04
Publisher
Unsupervised Learning

Transcript context

…ow this chain of thought in product uh then eventually you'd kind of have to train them, right? Like you'd have to train them for the same reasons you have to train like whatever models you should. Um And I just think that would be like >> not all want to know what the chain of thought our model has to get to a response for. >> Right. I mean, you know, I think I think it would be useful to some extent, and we are trying to capture most of that value, you know, either with like for summaries, uh which I think are kind of like a little bit of a stopgap. I think the longer-term solution here is having the model actually talk to you in real time, which you know, the later the latest version of Codex kind of do, the latest version of of the reasoning Gmail is kind of do, but I think I think that will get much better. Um Yeah, but but yeah, I I think there's something very exciting here about just like not uh not having the training signal fight against us, right? And not not Yes, because yeah, I I I think if you uh if you want to be able to understand what the model does in the long term, but you know, you're scaling a method that is like kind of going directly against it, it's you're probably not going to have a lot good time, right? At least the others have to learn a bitter lesson. Uh and so this decoupling, I think, is a very is an idea that gives me a lot of hope for our abilities to at least understand um you know, how these models' motivations and generalization evolve as they get better as they as they work for longer. Um yeah, I don't think it's a complete solution to AI safety and alignment by a long shot. I think it's just another tool in our in our toolbox. Uh but I am hopeful that I building our toolbox with technical tools like this, we can actually continue chipping away at the fundamental problems here. Yeah. Seems like almost like over the, you know, medium term, it's like something that's going to be incredibly helpful. Probably not the catch-all solution for for long-term alignment. Yeah, I mean, I think it's a tool that can help us understand I I I think it's actually very useful to like build understanding of long-term alignment, right? So for example, there has been this very exciting work um from um um um from a Panagiotis collaboration with Anthropic Labs uh on uh model scheming, where they investigate uh you know, depending on kind of what environment you you put the model in, how you train it, like is it is it prone to like start kind of like having hidden objectives that it pursues. And you know, what enables that that whole line of work is channel for monitoring, right? It's this notion of like, "Oh, you can actually inspect what the model's motivations are." So, you know, and I think from that, like, that might take us in a completely different direction in terms of mitigation, right? Like, maybe the right way is like changing the pre-training data of the model, or maybe it's something like, you know, the inoculation prompting from a topic. Like, I think I think those are very interesting ideas, but I think like having this ability to like understand this is very helpful to to evaluate these. Yeah, it's almost like rompting from a topic. Like, I think I think those are very interesting ideas, but I think like having this ability to like understand this is very helpful to to evaluate these. Yeah, it's almost like foundational for any further area of research. What are like the other research areas of then alignment that you're paying attention to or that you think are promising, you know, areas to focus on? Um Yeah, I think I think a lot of them a lot of the like longer term challenge with alignment is about generalization, right? Like, we can train our models to do well in in in in in our, you know, at least mostly to some extent. Like, we we we can mostly kind of control their behavior in the in the things that that, you know, are in distribution that that that that we trained for. Um But, you know, the things that are worrying some is like, "Well, what happens when the model's asked to do something very very different or it finds itself in a very different situation or it's like much smarter than it ever was before and and and, you know, it has this all these capabilities that's like we haven't really kind of thought about how to train for." And so, yeah, so so I think I think I think I think I think, you know, the study of like this kind of longer term value alignment is really a study of generalization. Like, what are the values that the model falls back on? Um Like, one line of research I'm very excited about here and something that we're uh investing in quite a bit is uh understanding like how that uh how the generalization falls back onto the pre-training data. Um Um Yeah, I I I and yeah, I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I guess over like, you know, the last 6 months, have your concerns around alignment increased, decreased? Like, how do you, you know, where are we kind of trending overall in in you know, with this work? I I I will to like the the the longer-term challenges of like follow your alignment, right? Or like what happens when you have various like models. The the way my thinking about the problem has evolved over the past few years is definitely kind of gone from you know, always this like very nebulous problem that like it's just like very hard to even grapple with or define uh to like, oh, you know, I I think we can actually make progress at it by very concrete technical solutions and technical insights. And this is why we've really been uh viewing LLMs like just a core part of of research and really uh you know, making sure that like we are, you know, designing our reasoning models, uh thinking about this, and we are, you know, and we are kind of like conducting our alignment research with like these reasoning models in mind and so forth. Um So, I think my general kind of uh belief that there's like a research path here that actually gets us to an extremely happy world uh has increased quite a lot. Um At the same time, right? I think uh my timelines to very capable models have definitely decreased a lot, right? I think we're we're not that far, right? And again, I don't think these models a lot. Um At the same time, right? I think uh my timelines to very capable models have definitely decreased a lot, right? I think we're we're not that far, right? And again, I don't think these models are smart in like any of the ways, but I think these models are just very transformative. And so, I'm quite optimistic like we can keep a good grip on like how we're doing on the alignment problem, how to roughly evaluate the risks of of of of of of our models or or the problems with them, but you know, but I do think we have to be, you know, as an industry as a you know, really prepared to like take trade-offs and you know, and possibly, you know, slow down development uh um depending on what we see. It's already interesting to see a lot of this work happening across the major labs. You know, the fact that you did this in collaboration with I think Anthropic and DeepMind and, you know, it seems like uh has that just come up organically or imagine like is there a lot of like alignment talk between, you know, the the major players, you know, uh given I guess the three of you are really at the forefront of all this? It's definitely some. I mean, there's definitely like shared interest in these topics, yeah. I want to shift a little bit to going inside of an AI. I feel like no no company probably the world has been more interested in over the last 2-3 years and you know, I think particularly what it's like to run a research organization. You know, we talked a little bit about this previously, but you talked before about how it's you know, important part of your job is giving researchers, you know, to to kind of have comfort and space to you know, almost be cave dwellers right and think about what the models will look like in a few years. You know, we're kind of alluding to it earlier. We're also in a time where it feels like there's just massive competitive race and you know, it's it's it's certainly you know, everyone's going really gangbusters on these coding models. I'm wondering like how do you actually operationalize this balance today and and you know, anything you kind of change in your thinking, you know, overseeing this organization around the right way to do this? You know, I focus on on just high quality experiments, organizing, you know, are we actually making progress, being honest with ourselves and you know, and promoting honesty about about the results. I don't think that has changed, right? And and you know, even though our work will evolve a lot, I believe we still have quite a lot of work left to do. And so I don't think it's like, oh you know, we need to wrap up all our projects you know, very very quickly. So yeah, I don't think there's fundamental change. I think what what does change is you know, a level of urgency to really kind of bring some of these things that we think are most promising at fruition. And then obviously you know, I feel like there's been you know, some very public internal moments of OpenAI over over years. You've been here for a long time. As you kind of reflect back like what were some of the difficult decisions that you guys made that maybe were like 51-49 that really you know, defined the company or…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence