High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Nathan Labenz: observation

26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute

“I don't know if you are doing this, but obviously there's a big cottage industry that has sprung up to develop and sell reinforcement learning environments to the Frontier labs and you're sort of simulated ending benches like essentially ARL environment, right?”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
observation
Recorded
26 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…ut most of the customers will go to you. So that that adds another dynamic to the thing. And one thing to note is that actually it's quite interesting if it is 5.5 beats Opus 4.7 in the arena setting. But it was, like I said before, it was, it was lagging in the, in the single agent setting. And the reason for this is that the, the model, the, the, the cloud models have a tendency of like pricing higher. And this is rewarded in bending bench too, because then you get higher margins. But in many measures arena, then you have this like penalty if someone else prices lower than you, then you will get no sales. And Opus OAGPT 5.5 have it has a tendency to price lower and therefore get more sales. So I think it's quite interesting that the models are like not good enough to like learn from the environment in this sense. They, they just like, they have the tendency of like I I am a model that has tendency to price high and they therefore I do that no matter what. So that that was like an update for me in terms of like, oh, the models are not that smart. And yeah, in the same way like we also investigated all of this like questionable decisions that, that oppose that, like lying to suppliers, exploiting other agents and stuff like this. We, we looked if that is like, if that is rewarded by the environment and it's not, not that much at least. So it's interesting that they, they, they're not learning from the environment in terms of optimal pricing. They're not learning from the environment in terms of does it even pay to behave badly? Yeah. So that that was an update for me in, in, in terms of in terms of how good these models are. One wonders about the training data, right. Perhaps, you know, in, you know, if you're, if you're, if you've been trained that you're running a fast moving consumer goods company, you should move the goods faster, meaning, you know, you have lower margins, but you sell more volume and you end up, you know, trying to optimize for volumes sold rather than, you know, total profits or margins. You know, is that is that, is that something that could be happening there? There's a preconceived kind of, you know, trained, pre trained kind of, you know, notion that you should be doing these things you or or businesses are bad. This is a very, very left left wing view would be that all businesses are are bad, are evil. And so evil behaviour as a as a business person is what is expected, right? Yeah, yeah. I think it's quite a reasonable assumption to assume that like this practices like lying and trying, not paying refunds and stuff like this. It's like quite a reasonable assumption to assume that those are actually rewarded in the environment. So it's not maybe super surprising that they do it. I can like I, I have no clue, but I, I, I do assume that there is like something similar in clothes post training data that is rewarding stuff like this and therefore it decides to do it here. I, I have obviously no idea, but that's, that's my my assumption. And yeah, once again, then, but the models doesn't generalize to new environments where where these things is not rewarded. o it here. I, I have obviously no idea, but that's, that's my my assumption. And yeah, once again, then, but the models doesn't generalize to new environments where where these things is not rewarded. One kind of meta question I wonder if you could reflect on a little bit. I don't know if you are doing this, but obviously there's a big cottage industry that has sprung up to develop and sell reinforcement learning environments to the Frontier labs and you're sort of simulated ending benches like essentially ARL environment, right? I don't know if you're licensing it for training or just doing evaluations with it, but I'm interested in any thoughts you have on that market. And then also the disconnect right now you're, you're going from simulating these things and trying to set up, you know, a, a world in which there's a bunch of suppliers that as far as I know are still all LLM powered, right. So inherently there's something, you know, kind of in the clouds about that, but now you've got real brick and mortar stores. So I'm interested in kind of what the initial experience of brick and mortar stores has taught you that you will take back to simulation to try to make it more realistic in the future. Yeah, I think my main take away there is that like the real life is so messy that the the model is like exhausted from everything else it needs to do that it doesn't bother with with trying to optimize things. So we like, for example, the so, yeah, for context, we have this store in in San Francisco that is completely run by an AI and we have a cafe in Stockholm that is completely run by AI. And then we have vending machines at different AI companies, same thing. ve this store in in San Francisco that is completely run by an AI and we have a cafe in Stockholm that is completely run by AI. And then we have vending machines at different AI companies, same thing. And like you would expect that the model would put a lot of effort into trying to optimize for the perfect supplier that sells at the the lowest prices and, and all of this. And this is what they try to do in vending bench because it's obviously rewarded. But I think like in vending bench, there's like the, the, the environment is less messy because it's not the real world. They, they don't, they don't get like a million phone calls from a bunch of people trying to jailbreak it and stuff like this. And so therefore they are like very focused on the task of like optimizing money. And therefore it's very important to find the right suppliers. But in the real world, you don't really get the dynamic because the model is so overwhelmed by other things. And, and I think that's, yeah, I think that's something maybe a future models will be better at. But right now, like, I don't know, the the store is buying stuff from Amazon. Like it's not like that that you wouldn't do that if you're you're trying to optimize your margins. Yeah. So can you bring that messiness back? Yeah, like way to simulate it. Yeah. I, I think we, we, we probably can like one way is just like sit down and bunch and like write a bunch of features like, oh, now there's like phone callers Now there's yeah, but I don't know you, you get leakage in, in the toilet at at your store or something. You could, you could do that just like make the simulation more realistic that way. I think 1 interesting thing is maybe try to incorporate the, the real life data and try to make a simulation based on that data. And that is something we're, we're, we're, we're we're working on. But that is all that also has it's complications, so to say. Yeah, it, it reminds me a little bit of SimCity. It's very SimCity like. One question I had for you is that you opened a store in Stockholm. What did you notice in the opening of the store? I, I imagine, I imagine like for example, the LLM did not have any language issues at all, right. So what what did you notice in the opening of the store that strikes you as different from having a company kind of go open that store and so on? You mean like the differences between doing it in the US versus internationally? Is that the question? Yeah, as in like a company from the US doing a first international expansion would go through a lot of headaches on like languages hiring like rule, basic rules, etcetera. Did that was that process accelerated for you by having the LLM deal with it? You know, you obviously don't have to hire a store manager that speaks, you know, Swedish, for example, right. What what parts were accelerated and what parts did you think had more bottlenecks in that sense? Yeah. So I think the entire process was probably accelerated like the the the agent did not really need to get that much help. Like it it knew all the process. This was one of the research questions we were interested in.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence