Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)
“You know, there's this annoying phrase like this is the worst the model will ever be.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 5 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…we might have an outer loop, which makes the system more open ended. But right now, my opinion, Genie 3, like all AI, gives you exactly what you asked for in the prompts and isn't creative on its own. Currently, the system only supports a single agent experience, but imagine how cool it would be if you could extend that to a multi agent system. Apparently, they are working on that. I mean, personally, I'm most excited about a new modality of interactive entertainment. You know, just imagine YouTube version 2. DeepMind sees the main use case of being able to train robotic simulations as being the real game changer. This seems plausible to me. I mean, like, the miracle of human cognition or in brains is that we have evolved to simulate the world without direct physical experience, which is expensive. This is basically the same idea. Right? Why train in the real world if we can just simulate any possible scenario in a computer, just like that Black Mirror episode? Here's a couple of examples they gave of using simulated environments to train an agent to do some specific language tasks. Now with Genie 2, they said they were happy if it was consistent even for 20 seconds. But now when you notice something inaccurate, it's very surprising. The key thing is that it now extends beyond the prediction horizon of the average human and the glitches are getting harder and harder to spot. They said that Genie 2 wasn't actually real time. You had to wait a few seconds between taking different actions. You know, it was low resolution, had limited memory. You know, it's a bit I mean, it was superficially really good, but it didn't look particularly photorealistic. Genie 3 changes all of that. So Genie supported around 10 seconds of generation, Genie 2 around 20 seconds. Genie 3 is able to simulate interactive environments for multiple minutes. This time around, they were a little bit more tight lipped around the architecture They wanted to focus on capabilities in the interview, and that's fair enough. I mean, it's understandable given that this is potentially a trillion dollar business, and Zuck will be sniffing around like a truffle hound. My my biggest concern with this is that as soon as Zuck gets wind of this, he is going to be getting out his checkbook. He's gonna go straight to Jack and Shlomi, and he's gonna be like, come on boys, $100,000,000, come and work for me. Zuck, mate, seriously, no. Don't do it. These guys, they they're doing God's work over here. You need to just let them let them do what they're doing. You can make it yourself if you want, Zuck. Leave them alone. I should say, I did joke at the end of the interview that if you are learning Unreal Engine right now, you might want to pivot to a different career. But the Google guys were quite grounded. They argued that this is a different type of technology. There are pros and cons, you know, which is fair. I should stress that as amazing as this technology is, it's still a neural network and it still has many important limitations. Certainly though, just imagine how easily you could generate an interactive motion graphics with this technology. echnology is, it's still a neural network and it still has many important limitations. Certainly though, just imagine how easily you could generate an interactive motion graphics with this technology. You know, that's something that Unreal Engine is meaning leaning hard towards in version 5.6. So do I need to fire my motion graphics designers, Victoria? Will users be able to use this? Not anytime soon. This is still a research prototype, and given the obvious safety concerns, they're gonna open this up progressively through their testing program. 1 question did come up in the press conference yesterday though, like, could it generate an ancient battle? And Shlomi said that it's not trained on that kind of data, wouldn't be able to do that yet. So I mean, certainly not a specific historical battle anyway. So it does sound like there are still some limitations. How can a system like this ever be fully reliable? Well, they did say that with better models, the trend is that they get more and more accurate, the glitches become fewer, and they expect to see further improvements. You know, there's this annoying phrase like this is the worst the model will ever be. But even as I said, they can generate some edge cases using a whole bunch of prompt augmentations, but it might just be turtles all the way down. And how do you come up with all of the rare black swan events that might happen? So what data was it trained on? They were quite cagey about this as well. It's probably safe to assume that it's been trained on all of YouTube and lots more besides that. How much compute does this thing need? Well, asked them that and they were a little bit vague about it. They said that it ran on their TPU network. So I'm inferring from that. It needs a crap ton of compute. However, I can say that it was demoed in front of me. It was very responsive. You put a prompt in. It thinks for about 3 seconds and then you're just in and it just works. They also mentioned some cool stuff about how, you know, like Genie can be used to train agents, as we said. But the agents themselves could be used to better train Genie 3, creating this virtuous cycle of iterative improvement. If you're in a world walking around say you go to cross the street, you sort of check the queues of the of the drivers, for example. Maybe there's not a crosswalk. And you need to know when to stop. You can see that they're slowing down, so that's when you would go. And the other agents should be simulated in that fashion.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.