High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Adam Marblestone: belief

30 Dec 2025 Dwarkesh Podcast Adam Marblestone — AI is missing something fundamental about the brain

“There might be other ways of learning, energy-based models or other things like that, that you can imagine that is involved in being able to do this and that the brain has that. But I think there’s a version of it where what the brain does is crappy versions of backprop to learn to predict through a few layers and that it’s kind of like a multimodal foundation model.”

— Adam Marblestone

Source trail

Everything needed to verify it.

Speaker
Adam Marblestone
Attribution
Verified speaker
Claim type
belief
Recorded
30 Dec 2025
Publisher
Dwarkesh Podcast

Transcript context

…Very naive question, but why can’t we achieve this omnidirectional inference by just training the model to not just map from a token to next token, but remove the masks in the training so it maps every token to every token, or come up with more labels between video and audio and text so that it’s forced to map one to each one? I mean, that may be the way. It’s not clear to me. Some people think that there’s a different way that it does probabilistic inference or a different learning algorithm that isn’t backprop. There might be other ways of learning, energy-based models or other things like that, that you can imagine that is involved in being able to do this and that the brain has that. But I think there’s a version of it where what the brain does is crappy versions of backprop to learn to predict through a few layers and that it’s kind of like a multimodal foundation model. LLMs are maybe just predicting the next token. But vision models maybe are trained in learning to fill in the blanks or reconstruct different pieces or combinations. But I think that it does it in an extremely flexible way. If you train a model to just fill in this blank at the center, okay, that’s great. But what if you didn’t train it to fill in this other blank over to the left? Then it doesn’t know how to do that. It’s not part of its repertoire of predictions that are amortized into the network. Whereas with a really powerful inference system, you could choose at test time, what is the subset of variables it needs to infer and which ones are clamped? Okay, two sub-questions. One, it makes you wonder whether the thing that is lacking in artificial neural networks is less about the reward function and more about the encoder or the embedding… Maybe the issue is that you’re not representing video and audio and text in the right latent abstraction such that they could intermingle and conflict. Maybe this is also related to why LLMs seem bad at drawing connections between different ideas. Are the ideas represented at a level of generality at which you could notice different connections?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence