Evidence receipt / evaluation
Published · transcript-backedGary Marcus: evaluation
24 Jun 2025 Machine Learning Street Talk Three Red Lines We're About to Cross Toward AGI (Daniel Kokotajlo, Gary Marcus, Dan Hendrycks)
“I take I take even, like, don't make illegal moves in chess to be a form of the alignment problem, like a very simple microcosm of the alignment problem, and they're still struggling with that.”
Source trail
Everything needed to verify it.
- Speaker
- Gary Marcus
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 Jun 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…There's there's still the the problem of if it's getting better, there's a question of is it able to do full automation, which is more of the key thing. And unfortunately for other abilities that probably the ability has most, well, its 2 main abilities would be crystallized intelligence, acquired knowledge, and readingwriting ability. But even for readingwriting ability, the models, when you ask them to write an essay, if you score them on the GRE score out of 6, they get like 4. 5 or so. They're not particularly great writers, and you can't automate writers that well if they're doing somewhat complicated writing. People still need to be doing that. And that's somewhat surprising because they've had so much data. They've had all the data in the world, basically. And likewise for coding, they have had GitHub. So that at least it's potentially the case that as you extrapolate them out, they'll get a lot of the low hanging fruit, But at crossing some sort of threshold for doing more automation or full automation, that could actually require resolving some of these other bottlenecks that are other cognitive ability bottlenecks that Gary's alluding to. I'm going to put a hold on this discussion. I think we did a good job. We're not going to convince each other, but we laid out the issues and that's great. I want to ask at least 1 other question because we have somewhat limited time. Do you think that we've made any progress on alignment? Do you think there's a conceivable solution to technical alignment? Are we close to it? Let's talk about the technical alignment problem. And I'll just put 1 thing out there, which is I think even though I don't think there's been as much progress as you do on capabilities, there's obviously been some. Like there's no argument there. I would say that a lot of it is sort of interpolation rather than extrapolation and we haven't solved the extrapolation problem. But interpolation, we've made huge progress mostly just in virtue of having more data and more compute. But, you know, there's no question that new systems are much better at interpolating, than previous systems, and that has lots of practical consequences for labor, for example. There's a whole bunch of things that you can use these systems to do that you couldn't use them before that's already affecting labor markets. Whereas my intuitive sense is on alignment, all we have is maybe a human reinforcement learning helps a little bit so that, like, if you ask these systems the most obvious question, like, how do I build a biological weapon? They'll decline. That's a little bit of progress. But, like, we all know that those are things that are easily jailbroken. Like my view, and you can agree or disagree or whatever, is like an alignment system still don't really do what we want them to do. There's still kind of sourcer's apprentice style problems. Are kind of problems of just like they don't quite fit, they don't do what we want. We have system prompts that say don't produce copyrighted material and they still do. Don't hallucinate, they still do. Like, they, you know, they can approximate what we ask for them, but, like, they don't really do what we ask them. I take I take even, like, don't make illegal moves in chess to be a form of the alignment problem, like a very simple microcosm of the alignment problem, and they're still struggling with that. That's my view. Why don't we start with you because you said this. Yeah. So there's some different notions of alignment or alignability, and I I I would distinguish between aligning proto super intelligences and aligning a sort of recursion that gives rise to super intelligence. Those are very qualitatively different. 1 is more modern model level and 1 is more process level. So I don't think you're going to solve that process level, 1 of doing a recursion and fully derisking that, anticipating all the unknowns and unknowns and solving a wicked problem in a nice clean way beforehand, which points to the necessity for resolving geopolitical competitive pressures and giving them an out to to proceed with a recursion more slowly or substantially forestall that. On the proto ASI things, I think that they follow the instruct instructions fairly reasonably. There are some other parts of…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.