High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“So there's like a really lot of unpack there here. So just like a little bit of the background, the reason that we started to create our own foundation on models like this realisation that what closed model providers are offering does not make sense for us economically.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…it. So I think that's like one of the big surprises. And for us, realizing that was like this big moment that validated something that we always strive for is to create an extremely efficient models. Because I think once you start to realize that the robot will need to create like the simulation 30 times a second, you just like realize the amount of tokens that is going to be burned for the simulations. So I think that was like one of the maybe exciting validations of the overall thesis in terms of architectural like a bunch of things that we can, I don't know, discuss in depth. We're planning to release on our mixture of expert architecture. Besides the dense models that we're already releasing, I think we finally were able to crack variable tokens architecture. It's also exciting and kind of teaches the model to invest more tokens where let's say the physics is challenging or like something necessitate to create more tokens. So anyhow, a ton of things are going on. We're gearing towards the release of our next module really soon. So yeah, busy times. Prakash asked where the real bottleneck is compute data or model design. Well, it's like our constraint. It obviously compute. We're like a company that funded the development of the model using profits from mobile content creation apps. So like we definitely compute constraints and like the big guys. And as to efficient inference, like recently it it really depends on the use cases, right? Like, so let's think about about a bunch of them. If you are, let's say you want to create like a real time avatars or like virtual environments, then OK, you can take like a huge model. You can do like a weight distillation to weigh kind of smaller architecture in terms of a parameter count, right? You can then they're like distillate, I don't know, two to four steps. And we're already at the point where for a lot of these use cases, we're like at the latency like way below a second, right? So I think we are hitting a point where these things are becoming production ready for some use cases. But for real time use cases, I think like avatars are extremely easy. We're going to see like a ton of avatars soon that they're going to be like, I don't know, like virtual teachers and visual customer support professionals etcetera to create an actual gaming environment environments. We still have a problem of having enough tokens for world consistency, right? So think about Genie free and the similar models, right? You typically create some kind of autoregressive model that has a lot of tokens that you already generated as your in in your kind of context window. And that blows up pretty quickly, right. So we were having some models that have like, I don't know, 30 seconds, like 60 seconds, It's still not enough to have an actual game. And if you think about that, the sort of brute force compression methods that we're using so far where for example, you just like sub sample tokens, they're like not really robust, right? Like just like imagine this scenario, what you, when you start to generate some kind of environment like my room, for example, and then I open the drawer and there's like a small coin there, right? t? Like just like imagine this scenario, what you, when you start to generate some kind of environment like my room, for example, and then I open the drawer and there's like a small coin there, right? You kind of expect that. Now when you get out of the room and you come back and you open the same drawer, you're still going to see the same coin right at the same place. But like this coin it just like this like tiny token that was generated and to creating a system that knows how to compress like the whole context in a way that's still going to preserve like these critical details. Well, we don't have it yet, right. So although like we do have like real time models that can do the things that the context is still missing there. And I don't think we're going to have like, you know, games that are running on the system like an actual games in the next quarter or two. And in terms of robotics, a lot of the use cases around robotics actually do not require like a ton of context window, right? Like think about like robotic arms and dexterity use cases. Then like the whole context is in front of you, right. You, let's say you want to, I don't know, figure out how the robot can create a sandwich. Well, everything is kind of the front of you. And then latency and auto regressive models. Yeah, like this part we already have. So you're going to start seeing them of robotic arms doing things like fairly quickly in the next quarter or two. Like so far if you're looking, a lot of these videos actually kind of speed them up, right. So it looks like the robot is something cool with its arms, but it's like, OK, extend the speed. I think that that's like mostly solved then the business question, why give a frontier model away? Yeah, so great question. So there's like a really lot of unpack there here. So just like a little bit of the background, the reason that we started to create our own foundation on models like this realisation that what closed model providers are offering does not make sense for us economically. OK, we were A at logics, we're a mobile creativity company. te our own foundation on models like this realisation that what closed model providers are offering does not make sense for us economically. OK, we were A at logics, we're a mobile creativity company. We really wanted to have AI models that are running, for example, on edge devices and where you don't spend an interest computer at all. And some point we realised that actually no one cares about creating a models like that. And when we try to see if we can work with closed model providers and serve it to our customer base, we just realise that it's completely prohibitive. And that's when we decided, OK, we are going to create an extremely efficient architectural and we can discuss like what's the bet there? But most of this boils down to the fact that you're creating an extremely compressive latent space. So videos represented by small amount of tokens and then you can add the top of it like a variable talking rate. Long story short, if really kind of go going to a close source providers, I think again I'll draw analogy to LLMS. Let's see kind of what happens there, right? Like Open AI and Entropic are trying to justify basically a trillion dollar valuation, right? And I think like the the story is kind of simple. If the tech is magical, it's hard to doubt it. So OK, if it's a magical tech, then we should put like a huge price tag on it. But when you're sometimes looking at the economical realities, it doesn't work out like that. There are like a ton of examples where the service is extremely valuable, but very, very hard to monetize, right? Like so now like we have like this interesting story where I think it's kind of clear that Chinese companies, you know, like deep sick moon shot, they're like really not that far behind in the lamps. But if you look at the evaluations of this companies at their last round, we're talking about, I don't know, 10s of billions, like maybe around like 50. No one is talking about the tree round. But like, wait a second guys, it's like the same underlying text. So what's going on? So I think it's with some of those internally calling like the CapEx trap. These guys spend like so much on the data centre, so much on compute, like such a crazy amount of money and creating such an expectations that they just like really try to create a business model that's a toll road, right? Like that every time that you touch their model, you're paying them. And maybe it could work in the past, but giving the the availability of Chinese model, I just like don't see how it's going to unfold like that. So imagine that in the world of world models, we are providing an alternative to people who do not want a toll road business model. So we're coming and say, listen guys, if you're not, if you're not hitting $10 million threshold, you can use the model for free. Just like you know, build something cool, get to some kind of traction and then we can discuss licensing. Once you're hitting 10 million of those revenue, let's discuss licensing. It can be multi year deal that's extremely predictable for you, so you can manage the cost etcetera.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence