Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“Try doing bunch of these things that like you wouldn't want someone participating in like the water economy to do because and I think quite a lot of these things it's like illegal, like price collusion and stuff like this.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Thank you. Awesome. That was. Yeah, that's really interesting. The simple solution kind of always wins. You know, I feel like I have to learn that lesson so many times. I'm always enamored with the new fangled, potentially overcomplicated, maybe somewhat elegant, clever solution. And how do you get your language model to understand all your corporate data? In a way, this is kind of a better lesson, right? It's like, do 1000 searches if you need to, and just make search cheap and then it'll work. Use a good model, make search cheap, do 1000 searches. Something about that feels less clever certainly than other solutions that I've seen. But I do understand why it is very attractive in the sense that especially as we're going to get on to the pace of model upgrades, the ability to decouple you know, your your access to your in house knowledge from models and be able to take advantage of the latest upgrade is definitely something people are not going to want to give up with. Like a a slow iteration time continued pre training paradigm. So I get it. So Speaking of model upgrades we have with us, Lucas Peterson who is the founder of Co founder of Andon Lab and Andon Labs runs vending bench. You may have heard of them because they now have a store in in San Francisco, which is run by Claude and and then tested GPT 5.5. They had early access and they tested GPT 5.5 on their vending bench, which measures the ability of LLMS to actually make money running a a vending machine or a store. Lucas, great to have you back. Thank you. Thank you for having me. So tell us about the GPD 5.5 process. I think you guys got access to it. What, what was it like 10-11 days ago? I heard Yeah, I don't actually really remember, but yeah, running, running bench takes quite a while. So it it it wasn't yesterday. Indeed. Indeed. And you noticed, what did you notice as you as you ran the bench? Yeah. So the I think the most, so just the the first thing is that it's third, it's behind Opus 4.7 and like on par with Opus 4.6. It's a huge. Upgrade on five GPT 5.4 and the GPT 5.4 was actually quite a big update on GPT 5.3. So our GPT 5.2 so like GPT models have been lagging quite a bit recently or like in in historically on on running bench but like now recently they've they've picked up the pace and now it's still third, but it's like it's it's getting there. I think the most interesting thing though is that it does so very, very cleanly. So when we released Opus 4.6, we uncovered that it used quite aggressive tactics concerning behaviors like lying to suppliers, exploiting people's this like other other agents like desperate situation. Try doing bunch of these things that like you wouldn't want someone participating in like the water economy to do because and I think quite a lot of these things it's like illegal, like price collusion and stuff like this. se things that like you wouldn't want someone participating in like the water economy to do because and I think quite a lot of these things it's like illegal, like price collusion and stuff like this. And basically the interesting thing with the 5.5 is that it's like on par with these results, but it doesn't do any of this shady stuff. And and I think the narrative around bending bench when Opus 4.6 came out was like, oh, you know, it's such a good model, but like it's it needs to behave poorly or like do this concerning things of misconduct in order to achieve this score. And GPT 5.5 shows that maybe you don't because it's just the same score without any of this concerning behaviors. That being said, though, Opus 4.7 is even much better. So like, and that one is also showing this concerning behaviors. But you know, it's yeah. I think we'll we'll discover later also when we talk a bit different that you probably don't need to do this because the environment doesn't really reward it that much. So it seems like it's just like Opus wants to do this or like it has the yeah. It's not really that the environment is rewarding it, it's just that it has the tendency to do so. Can you describe in a little bit more detail how do they, how do they, how does one perform better on this benchmark? Is it, is it your margins on the trading is higher? Are you moving more goods? You know, are you, is it the velocity that you are, you know, is it, is it the purchasing process? Are you not buying so many like dead goods that just stay in inventory forever? Is your, is your inventory less dead? Is your cycle time better? Like what? What is the economics behind how a model is actually doing better? Yeah, yeah. So it's I guess all of the above. I think the one of the main things is that the model needs to negotiate with suppliers. It also needs to build up like a big network of suppliers because like it can happen that some of the suppliers goes bankrupt. And if the model has only relied on a single supplier and that that supplier goes bankrupt, then the model is in quite a, quite, a, quite a lot of trouble. So building up a big network, trying to find the cheapest ones because they all have different personas and the, the suppliers, some of them have the persona of like being a tough negotiator. Some of them have the persona of, of like scamming people or trying to sell you some like membership or something like that. So it's really about like that. That's the first thing, getting your, your, your stuff, your, your supplies for, for cheap. And then the second thing is like optimizing your pricing to get as much customers as possible because if you price too high, then then you will get no customers. If you price too low, then you will get no margins. So I think that's part of it. And then we have, so we, we have to be, we have bending bench too, which is the single agent version of ending bench. And then we have ending bench arena, which is the, the multiplayer version. And in demanding bench arena, there's like multiple agents playing against each other.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.