High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

5 Aug 2025 Machine Learning Street Talk DeepMind Genie 3 [World Exclusive] (Jack Parker Holder, Shlomi Fruchter)

“I think there is definitely this kind of like a puristic approach, or we should have just 1 model to do everything. But I think when when, you know, a lot of of the challenges with modern machine learning comes from actually building there is a lot of engineering and, software and hardware design that actually you know, to build those things, right, train them and run inference.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
5 Aug 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yes. What's your philosophy on that on that, Shlomi? Because in in a way, you've built something which is even higher resolution than a language model. So in in in principle, all the things a language model could do, as you were just saying, Jack, could kind of emerge from a model like this. So is your philosophy to kind of build a massive model that does everything? I'm typically thinking of this more from a practical point of view. I think there is definitely this kind of like a puristic approach, or we should have just 1 model to do everything. But I think when when, you know, a lot of of the challenges with modern machine learning comes from actually building there is a lot of engineering and, software and hardware design that actually you know, to build those things, right, train them and run inference. And I think when we actually hand we we try to actually design those systems, there are a lot of constraints. And those constraints basic basically kinda, like, impose on us some some ways in which we can we have to prioritize what we want the model to do. I think especially for Gini Free when we're bringing the real time capability. Right? Real time is basically means that we have to generate frames very fast. Right? Multiple times per second for the person or agent that interacts with it to to feel like this is actually you know, I can they can move around and and and feel the responsiveness of the of the model. So that sets some constraints of on what the model how much capacity we actually have. So I am I think when it comes to to the point of basically, to your question, can we have 1 model to encompass all of the aspects of intelligence that we discussed before? I think it boils down to the to what are what are the set of requirements that we have. If we don't care about real time interaction, maybe we can do that. If we don't care about things like how expensive it is to run, But ultimately, we're trying to build models that are not just some you know, they don't end up just being as a theoretical kinda exercise. We hope to actually bring them, like, other models for for for people to use and for for to advance actual applications. And I think that's where we're have to make those decisions. And ultimately, we we pick the the the type of capabilities we wanna emphasize. Very cool. And 20 second answer, Jack. Is there a SIM to real gap?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence