High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Tim Scarfe: evaluation

5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)

“Perhaps in the future, we might have an outer loop, which makes the system more open ended. But right now, my opinion, Genie 3, like all AI, gives you exactly what you asked for in the prompts and isn't creative on its own.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
evaluation
Recorded
5 Aug 2025
Publisher
Machine Learning Street Talk

Transcript context

…Exactly. Yes. Someone from our team is actually playing this. They're pressing the w key to move forwards. And then from that point onwards, every subsequent frame is generated by the AI. Around the same time last year, you'll probably remember this by the way, DeepMind's Israel team led by Shlomi Fruppter showed diffusion models simulating the Doom engine. The system was called Game Engine. It's almost a meme at this point how Doom runs on, you know, calculators and toasters. But here was a neural network confabulating a Doom game frame by frame in real time. Like, look at how it just knows what the health is. Like, you can shoot characters. You can open doors and navigate around maps. You know, occasionally, was slightly glitchy. But this is just unreal. You know, you could just simulate Doom at 25 frames a second on a single TPU. The only limitation, of course, was that it could only do doom and nothing else. So last week, we waltzed our way into London, and Jack and Shlomi gave us a demo of Genie 3. Honestly, I couldn't believe what I was seeing. The resolution is now 7 20 p, which is firmly in the good enough territory to suspend disbelief. It's real time. It can simulate real world photorealistic experiences, which can continue for several minutes before running out of context. Shlomi had his hands all over v o 3 by the way, and they seem to have combined elements of the Genie architecture with v o producing something I can only describe as v o on steroids. Unlike Genie 1 and 2, the input is now a text prompt, not an image, which they argued is a good thing from a flexibility perspective. But it does mean that you can no longer take a photo of a real place and generate from there. 1 of the main features of Genie 3 is that it has a diversity of environments, a long horizon, and promptable world events. Now on the world events, let's take this ski slope example. We might type in another skier appears wearing a Genie 3 t shirt or a deer runs down the slope. And there you are. Things just happen in the world. They say that this might be very helpful for modeling things like self driving cars where you can simulate rare events. But I was left thinking that this is just turtles all the way down. How can we write a process to prompt the potentially infinite number of rare things which could happen in a scene? There was an example they showed of flying around a lake and it was amazing. But I was like thinking, where are the birds, mate? Like, can can you can you type the birds into the prompt? The team believes that we haven't yet had the Move 37 moment for embodied agents, you know, where an agent discovers a novel real world strategy. They see Genie 3 as the key to enabling that, but the real world constantly surprises us because the real world is creative. Creativity simply means that the tree of things which can happen keeps growing. New branches and leaves just keep appearing. Perhaps in the future, we might have an outer loop, which makes the system more open ended. But right now, my opinion, Genie 3, like all AI, gives you exactly what you asked for in the prompts and isn't creative on its own. we might have an outer loop, which makes the system more open ended. But right now, my opinion, Genie 3, like all AI, gives you exactly what you asked for in the prompts and isn't creative on its own. Currently, the system only supports a single agent experience, but imagine how cool it would be if you could extend that to a multi agent system. Apparently, they are working on that. I mean, personally, I'm most excited about a new modality of interactive entertainment. You know, just imagine YouTube version 2. DeepMind sees the main use case of being able to train robotic simulations as being the real game changer. This seems plausible to me. I mean, like, the miracle of human cognition or in brains is that we have evolved to simulate the world without direct physical experience, which is expensive. This is basically the same idea. Right? Why train in the real world if we can just simulate any possible scenario in a computer, just like that Black Mirror episode? Here's a couple of examples they gave of using simulated environments to train an agent to do some specific language tasks. Now with Genie 2, they said they were happy if it was consistent even for 20 seconds. But now when you notice something inaccurate, it's very surprising. The key thing is that it now extends beyond the prediction horizon of the average human and the glitches are getting harder and harder to spot. They said that Genie 2 wasn't actually real time. You had to wait a few seconds between taking different actions. You know, it was low resolution, had limited memory. You know, it's a bit I mean, it was superficially really good, but it didn't look particularly photorealistic. Genie 3 changes all of that. So Genie supported around 10 seconds of generation, Genie 2 around 20 seconds. Genie 3 is able to simulate interactive environments for multiple minutes. This time around, they were a little bit more tight lipped around the architecture They wanted to focus on capabilities in the interview, and that's fair enough. I mean, it's understandable given that this is potentially a trillion dollar business, and Zuck will be sniffing around like a truffle hound. My my biggest concern with this is that as soon as Zuck gets wind of this, he is going to be getting out his checkbook. He's gonna go straight to Jack and Shlomi, and he's gonna be like, come on boys, $100,000,000, come and work for me. Zuck, mate, seriously, no. Don't do it. These guys, they they're doing God's work over here. You need to just let them let them do what they're doing. You can make it yourself if you want, Zuck. Leave them alone. I should say, I did joke at the end of the interview that if you are learning Unreal Engine right now, you might want to pivot to a different career. But the Google guys were quite grounded. They argued that this is a different type of technology. There are pros and cons, you know, which is fair. I should stress that as amazing as this technology is, it's still a neural network and it still has many important limitations. Certainly though, just imagine how easily you could generate an interactive motion graphics with this technology.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence