High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Nathan Labenz: recommendation

13 Jun 2026 The Cognitive Revolution AI in the AM — Week 2 Highlights (June 2026)

“It can only do probably the frog game at the end of this training. But building out a world where we have these little role-specific AIs doing their jobs, doing it really well, I think that creates a much more buffered environment that's probably a lot more resilient to another generation of AI that's just like amazing at everything coming in and kind of shocking the system in such a profound way.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
recommendation
Recorded
13 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…But what Fable ended up doing was finding satellite images for this area. And that's how you get these colors and that's how you get the textures. But then to make it to scale and to make it accurate, it actually fetched elevation data from NASA. And it combined those two to sort of make this to scale. And that is what blew my mind, right? Because usually when you vipe coding stuff, you give an end objective and this objective is vague and there are 100 steps in the middle where humans would take decisions differently. And usually vipe coding doesn't work out very well because the quality of the decisions the models make aren't always great. But Fable made such high quality decisions where it eventually ended up creating something that exceeded the expectations of what was initially a very vague objective and did so in really smart ways. So I'll give you another example, right? So you see all of these trees and V1 of this project did not have any trees. And I was like, hey, I think we're missing some trees here. I would love to add them. And I would have been completely okay with just randomly creating these trees. But what it actually did was it analyzed the pixels on the satellite images. It found out the ones that could potentially have trees, so the ones that were green maybe, and added trees only on those spots. But it didn't stop there, right? It realized that because it was analyzing pixels, some of those pixels were white. So you can see that there is snow in the mountains far ahead, and it also added snow. So it just exceeds your expectations in these small and subtle ways, makes really smart decisions. It's like having a really, really smart employee with extremely high agency who blows your mind every single time. Friday morning brought the week's cleanest empirical result on the recursive question. And it's not from a lab. Here's one other thing I'll just touch on real briefly. This is thoughtful. This is a company started in part by a woman named Karina Wen, who used to be at Anthropic, then she was at OpenAI, now she's doing this. And this is maybe one of the more telling, you know, it's kind of vibes, it's kind of quantitative, it's a very idiosyncratic task, but it's also a very relevant task to the future. Can you get your top model to train a small model effectively to do a job for you? And this is something that, as you can just see with these bar graphs here, the particular frogs game thing, it's kind of like a Sudoku type puzzle that they're training a small model to do. And the big models can often just solve it, but the small models can't. So the challenge for the big models, can you train the small model to solve it? And this involves all these little tips and, you know, not tips, but tricks and know-how and kind of, you know, hard lessons learned by post-trainers who've been in the trenches doing this. the models up until Fable basically didn't really move the needle on what the small models can do. They basically just couldn't do this sort of post-training effectively. But here we see more than 10x improvement on small models' ability to do these tasks. And again, I think this is one way that it could be really good, right? If you had like very narrow, very small, very role-specific small specialist models in all these different niches, That could be a great world, right, that gives us a lot of abundance in a very affordable way. In this little small model that got post-trained to play the frog game, it's not going to go out of control, right? It is small. It can only do probably the frog game at the end of this training. But building out a world where we have these little role-specific AIs doing their jobs, doing it really well, I think that creates a much more buffered environment that's probably a lot more resilient to another generation of AI that's just like amazing at everything coming in and kind of shocking the system in such a profound way. So I think this goes to show, again, just wow, what capabilities we have that we have not absorbed and gives them a little bit of a foreshadowing of what a world of tons of small but highly performant AIs could look like in all these different little niches and how we get there, right? There's not enough human post trainers, but now we have Fable to do the post training. So watch that space. Now let me introduce a voice you'll hear a few times this episode. Prince, an anonymous practicing lawyer who built Prince Bench, a legal reasoning benchmark the labs themselves watch. He guards the anonymity, so you'll hear him and you won't see him. On Friday, he gave us a close reading of Anthropic's own launch documents that I haven't seen anyone else do. labs themselves watch. He guards the anonymity, so you'll hear him and you won't see him. On Friday, he gave us a close reading of Anthropic's own launch documents that I haven't seen anyone else do. So give us some alpha that you have picked up this week. This could be from your own testing. It could be from the system card. I know you're often a close reader. So we're looking for like the deep cuts of the things that you think even the AI obsessed have. overlooked or not fully appreciated yet?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence