Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)
“But in Genie 2, there was an ST transformer, so a special temporal transformer, which was conceptually quite similar to like a VIT. And there was a latent action model, which means even from non interactive data, you could infer some low cardinality action space.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 5 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. So I think you what you said about Veo or more generally about video models is right. Like, there is a way we can think about them as somewhat a world model, but it's not really it doesn't allow us to actually navigate or interact with it completely interactively. And I think that's that's 1 of the limitations of of video models that with Genie free we're trying to address. And basically, in the examples that you've seen, we're able to because Genie free generates the experience and what's the what we see frame by frame, it lets the user or the agent that is using it basically control where it wants to go in every like, in a very low latency. That allows basically exploring the environment and and creating new trajectories that are not predefined like video models. So in the examples that you've seen, for example, you can see the character or the agent in this video moving around, maybe going back to the place they're already been in before, and everything remains consistent. And I think that's a very remarkable property or capability of the model, the ability to preserve and the consistency of the environment along very long trajectories. Yes. And and even Genie 2 had some kind of object permanence and consistency, but nowhere near as much as we have now. But we'll come back to that in a second. We can't say too much about the architecture for Genie 3. But in Genie 2, there was an ST transformer, so a special temporal transformer, which was conceptually quite similar to like a VIT. And there was a latent action model, which means even from non interactive data, you could infer some low cardinality action space. And then those went into a dynamics model. I think what we can say about the architecture that might be interesting is that definitely because of the interactive nature of the problem or the setup, then the model is autoregressive. So what it means that it means that the model generates frame by frame and has to refer back to everything that happened before. Right? So if, for example, we are walking around some auditorium or some other environment, basically, if revisit a place that we've already been to, the model has to look back and and understand that this information has to be consistent with what's happening in in the next frame. So I think the interesting the interesting point here is that everything here like, the consistency is emergent. There is nothing explicit. The more the model doesn't create any explicit 3 d representation. And it's unlike, you know, methods other methods like NURBS and Gaussian splatting. So I think that's that emergence kind of like capability is very interesting and surprising for us. Yes.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.