High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)

“I think what we can say about the architecture that might be interesting is that definitely because of the interactive nature of the problem or the setup, then the model is autoregressive.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
5 Aug 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yes. And and even Genie 2 had some kind of object permanence and consistency, but nowhere near as much as we have now. But we'll come back to that in a second. We can't say too much about the architecture for Genie 3. But in Genie 2, there was an ST transformer, so a special temporal transformer, which was conceptually quite similar to like a VIT. And there was a latent action model, which means even from non interactive data, you could infer some low cardinality action space. And then those went into a dynamics model. I think what we can say about the architecture that might be interesting is that definitely because of the interactive nature of the problem or the setup, then the model is autoregressive. So what it means that it means that the model generates frame by frame and has to refer back to everything that happened before. Right? So if, for example, we are walking around some auditorium or some other environment, basically, if revisit a place that we've already been to, the model has to look back and and understand that this information has to be consistent with what's happening in in the next frame. So I think the interesting the interesting point here is that everything here like, the consistency is emergent. There is nothing explicit. The more the model doesn't create any explicit 3 d representation. And it's unlike, you know, methods other methods like NURBS and Gaussian splatting. So I think that's that emergence kind of like capability is very interesting and surprising for us. Yes. And and even Genie 2 had emerging capabilities like parallax, and, you know, it could model certain forms of lighting and and so on. But this just blows my mind. Yeah. I Sean, well, you're you're involved in that Doom simulation last year, and even that just blows my mind. Right? So we all played Doom in 1993. It was 1 of John Carmack's finest. And now you're saying that I mean, the work that that you folks did last year. You've got a neural network model which is sub symbolic. So there's no explicit model of the world. You don't know where the doors are. You don't know where the lakes are, where the maps are, and so on. You just kind of take a, you know, a sample, a traversal through this space, and and it just produces the game in pixel space. I mean, that's yeah.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence