Evidence receipt / preference
Published · transcript-backedTim Scarfe: preference
5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)
“You know, I'm a huge fan of open endedness, for example. And certainly at the moment, when we prompt models, if we're quite generic in what we put in the prompt, then we tend to get quite simplistic answers.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 5 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So I think it's still early to say exactly how word models like GenieFree will be actually used for AI research. I think we can only kind of directionally say. We I I in general, I think we we still we see it in also in other generative models that there are some capabilities that we actually discover. Right? And we don't necessarily know that they're there. And then through the interaction development, we're actually seeing them emerge. For example, you know, with with Vell, we just recently, like, few days ago, we them we we kinda, like, shared that you can write, like, some text on a on a photo and provide it to Vowel, and it just, like, it's it's it reads the the the text, and it follows also the spatial instructions. Right? And I think that's that's, for example, something that we didn't necessarily explicitly train them all to do, but it it's capable of doing. And I think here as well, the capabilities of Gini Free that we're exploring are still we still discover new things. And I think that's something that we hope that by first, by having more, like, you know, testers and external testers that we already shared some some like, we basically previewed the model to and give us feedback. So we hope that through this kind of engagement with the community, and we can better see how those models will be useful. And that's something that I expect to take some time as we we basically try and understand the best application. You know, I'm a huge fan of open endedness, for example. And certainly at the moment, when we prompt models, if we're quite generic in what we put in the prompt, then we tend to get quite simplistic answers. You know, a lot of people doing computer graphics when they prompt image models, they they have so much specificity and they deliberately take it, you know, on onto the tail of the distribution so they get something that's novel and interesting and so on. And and the real world just always produces a sequence of artifacts which are novel and interesting. You know, like, you get random NPCs walk onto the onto the screen and cars go in and and so on. And is my intuition correct that that at the moment, as good as it as it is with with Genie 3, you you tend to get quite a specific scene. And you don't have like random kind of planes flying over and sort of, you know, just random things happening. Yes. That's a really good intuition. Right? So it definitely is the case that the model is very it's very aligned with the text prompts that that it's given. So therefore, there is a lot of emphasis placed on the quality of the text prompt to describe the scene. But I I actually wouldn't see that as an imitation. I would see it a strength. Right? So firstly, it means that there's actually a lot of human scale still involved to create really cool worlds. Right? And you see some of the examples we showed you. We have some very talented people that can do amazing things with these models. Right? And and there is actually a lot of value add there to do that. Right? So actually, it's a tool that can really amplify already creative humans in new ways. And I'm definitely not the best at doing this. Right? And I can tell you that it does it is really impressive when someone is able to do that. But on the on the flip side from the the agent perspective as well, right? So when we're talking about designing environments for agents, and and you referenced POET, which was for me like POET and dwell models were the 2 papers that I just thought were eventually on a collision course. Right? And that's basically when I started my research career. And I think POET was fundamentally limited because of the the environment encoding being an 8 dimensional vector, but also the fact that there was not really any notion of interestingness as well. Right? And in your recent interview with with Jeff, right, he's obviously talked about how this problem is largely now solved with foundation models. Right? So these foundation models can not only define what's interesting based on standing on the shoulders of of human knowledge. Right? But they can also steer the generation of worlds in things like omnipic to do this. And that's in that case, it's done through code. But here, we have text as a substrate as well. So in theory, this these kind of open end algorithms that use language could actually be quite strong places to have these kind of, like, notions of interestingness and agents steer tasks through that space as well.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.