Evidence receipt / uncertainty
Published · transcript-backedNathan Labenz: uncertainty
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“There's never any like really like step change for like now you're in a new environment or anything like that. It's just like a continuous loop, but whenever it hits some kind of token threshold, which will change every day, maybe it's 100K today, I don't know.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…sure you could if you try hard enough. Nothing like daily token when you said you can't read everything. It just occurs to me that like, and especially you talk about getting like tons of phone calls. What is the daily token budget in either millions of tokens or dollars or both that it actually costs to run the store? I'm kind of curious as to how the AI manager compares to a human manager in terms of just, you know, cost to have somebody do this job. Yeah, I, I should know these numbers, but I kind like I, I don't I, I think it's something like maybe maybe $100 per day or something for like maybe both stores, but I'm, I think it might be less, I don't know, some that's order of magnitude I think. Well, that's definitely notably cheaper than humid. Sounds like still distinctly worse performance though. So we're, we're staying for a minute. I mean, like the six months ago the vending machines were like OK, but not that great. One year ago they were quite horrible. So like within one year we went from like they can't do anything to now like vending machines are too easy. A store is visible like 6 months from now, probably a store will be too easy as well. I don't know. And it would be interesting what you could do then. And you think the main difference is going to be this sort of meta cognitive type stuff. It's not like what I'm hearing you say is it's maybe not anyone micro task that it's unable to do, but it's more you described it as exhausted. It's kind of it's failing to zoom out and take stock of its situation and kind of say how could I be doing better here overall? Is that the big frontier that you think? And you know, certainly that kind of seems highly related to getting a IS to do AI R&D more effectively as well, right? They, they can already write the code, they can already monitor the logs. But can they do that zoom out and kind of something like taste of, you know, what should I really do next to be most effective in the big picture? It seems like it's kind of the same frontier for both of these seemingly like quite different occupations that a IS might soon be playing. Yeah, I, I do agree. And I think that's partly why we're doing this. Like I think AIRND like lost control from like autonomous replication that is quite scary. And, and I, I hope that we can provide some valuable insight into that, even though we're not like tackle it, tackling it heads on. I think, I think most of the things that we're measuring here like translates to to those scenarios as well. And yeah, like, like you said, like being overwhelmed by a lot of data and a lot of context memory issues, stuff like this is, is definitely, definitely one of the things that is is lacking on a meta level right now. So one of the questions I had for you is how does your harness look like? Because you have this context length, right? The models have context length and then you have some tool calls. And when you say exhausted, is it a function of the context length where the model kind of only kind of recognizes like, you know, the last 100,000 tokens or whatever and the rest of the million token window is kind of, you know, not parsed properly. e context length where the model kind of only kind of recognizes like, you know, the last 100,000 tokens or whatever and the rest of the million token window is kind of, you know, not parsed properly. How does your compaction work? I imagine over the course of the vending bench you hit limits or either in terms of, you know, whatever limit that you set for the context window. So how do you kind of, is it end of day, kind of you do a compaction in order to start the next day and then you have a, you know, you restart the context window. So when it boots up again, it's like, OK, I'm on day five and this is my starting position in inventory. This is my starting position in like in, in cash. These are the outstanding orders which haven't come in, etcetera, etcetera. How does that work? How does your harness work? Yeah, it's, it's by the sign, extremely simple. Like we, we the sign is simple because I have too many friends who make some complicated, complicated harness. And then the next, the next model release, they have to throw it all out because the new model just works without it. So it's very simple. Like it's, it's just like it has it's a continuous loop. There's never any like really like step change for like now you're in a new environment or anything like that. It's just like a continuous loop, but whenever it hits some kind of token threshold, which will change every day, maybe it's 100K today, I don't know. But somewhere we're experimenting with it. We're compacting the the the thing and then it's like starts to build up a new, a new context for like from caching reasons. You don't have a sliding window, all of this basic stuff. imenting with it. We're compacting the the the thing and then it's like starts to build up a new, a new context for like from caching reasons. You don't have a sliding window, all of this basic stuff. But it's yeah, it's a basic thing with a bunch of like sub agents for a specific task like browsing and stuff like that. Yeah. Anything else interesting to say there? Yeah, but I, I think like the main thing is it's, it's very simple by design because we want to, we, we think that the, the, the, the better the models get, the, the simpler the harness will be. And we want to like surf the frontier. I'm sure we can like, I don't know, make like a vending machine harness and and like get some percentage better performance if we do that. But that's not really the point of what we're doing. Have you tried testing things like open claw? I mean, that's obviously not the simplest available harness, but it is something that has a lot of market penetration, right? So I'm kind of wondering if, and it would be simple for you to implement and upgrade on an ongoing basis. How do you think about kind of, you know, Lucas's simple harness versus the simplest thing that's like toward the frontier that you could easily install? Yeah, I think I think our thing is quite similar to open claw. Like we, we, we've been working on it for for quite some time, like long before open clock came out. But and, and there's a bunch of things that are a bit like basically like most of our time goes into goes into like the integration and stuff. And, and I think all of that you would still need to do with an open claw. We could, I guess, replace our agent loop, but we also, we want to keep it simple because we have like more control and we, I think it's, it's like a more accurate measure of, of we're different tier of the AI models are. And we're like more interested in measuring that than trying to push the, the performance because like in the future, the models will be smarter than humans and probably like a good scaffold will not help the models. So yeah, that's, that's the reason. But we could like, that is something we could do. It's just like when we started open cloud wasn't the thing. So we, I guess we built our own open cloud before it was called open cloud. But but yeah, that's the reason. What do you think happens next? So you have the models are now producing profit, right? The, the, the, the, the stores, the vending machines are now profitable, correct? Yep. And do you think there is, you know, on, on the last time you were on the show, we talked about where the ceiling is. So what, what do you think happens next? You know, in terms of the retail store, what do you expect for the next leap in the model? Like just just to get a calibration so that we can see if it's a linear or exponential and the next model lands. What like what do you expect in the in the next version? Yeah, I, I think it's quite hard to measure improvements on this like live, live real life deployments because you don't have AB test, you only have N = 1 and stuff like this. So I don't think you would like see a step change once a new model comes out.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.