Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“Overall, Pangram is quite accurate and yet we have at least one example out of 400 or so essays where I think the zero score I would confidently assert is wrong and unfair and should not be the basis for like a pylon.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…th content that you thought would have meaning. And it turns out neither the the author writing it didn't just didn't care enough to actually write anything meaningful. And it's literally slop, right? But on the other hand, if you have a writer who actually had like original thinking and actually cared about, you know, what they were thinking about, but then they used the AI to express themselves, does it actually matter that much? Then my report roughly 400 podcast intro assets through Pengram and a close look at the four it flagged as AI. It calls to mind my reaction to Fable where I was just kind of like, I don't think I should be so precious anymore. I, you know, I need to figure out some sort of merged way of working some hybrid output, you know, should probably be the norm now. I think you're heuristic of like if something drew me in and in the end I feel like I wasted my time. It's kind of like a time well spent metric from Facebook from back in the day, right? It's if I'm spending time trying to make sense of something that in the end I feel icky about, then that's clearly a problem. I guess just to close the loop on the, you know, what can we say about Pangram Labs based on this experiment? But I'll give you a rendering on the four that it said we're entirely AI. Two were entirely AI and admittedly so this one I think is, I guess I would come down and say, fair enough. I started with this. If you're only listening on the audio, you can't see this. But I made 10 different edits over the course of 10 minutes that cleaned the thing up. Just because something got a 0% on Pangram doesn't mean that it was like uncritical or that there was no human, no meaningful human role in the authorship. And then this next one, by the way, actually goes a lot further. So if I you can see just the time that this took, I am from 4:28 PM all the way through 5:21 PM. So more than 50 minutes continually seem, I mean, you can't, you know, I never tabbed over to anything else, but pretty consistently focused on this document making changes and it gave me still a 0. But yes, I think, you know, what can we say? Overall, Pangram is quite accurate and yet we have at least one example out of 400 or so essays where I think the zero score I would confidently assert is wrong and unfair and should not be the basis for like a pylon. You know it would the the crowd, the digital mob would be like in the wrong for piling on somebody, for passing off my snowflake intro SA or, you know, for attacking it as being AAI slop output. I think it I can show this edit history and everybody should agree that like, yeah, you put in an hour, you basically rewrote almost every section and somehow Pangram still gave you a 0. From this I would say you cannot convict in a, you know, reasonable doubt system purely based on this sort of thing. And yet at the same time you can pretty much trust the Pangram signal as a consumer. I think you can trust it as a judge. I think you should be more cautious. Also from that Thursday, the day Claude Fable 5 came back online, Palantir's Alex Karp had spent the week telling companies the Frontier Labs will absorb their workflows and steal their IP. Prakash's response in brief. You look at the Fortune 500 Nike. What does Nike have to fear from Anthropic? You get all the physical businesses out of the way and what you're left with is the pure IP businesses, right, software production, maybe pharma. I'm not I'm not so sure the paperwork businesses, banking, paperwork compliance businesses, accounting, tax compliance, regulatory, all of these things which are paperwork businesses, right? Those are all of the businesses where you have IP or relationships built up over years where if you have an Anthropic go in and they read through your entire workflow and processing, they can basically absorb all of that into the model. So that is where I think the risk is. The Frontier Labs are also horrible at sales, right? They're not. You look at IBMIBM has like 70% of the staff are basically sales engineers and the sales engineers are there to basically help you implement, maintain, you know, do all the grunt work, etcetera. The labs are not doing that. The labs have decided, especially Anthropic has decided to do this very lean structure of having almost no people at all and just putting out the models. And then just saying to like these enterprise teams here, you can go ahead and use it or not use it. And I'm the CTO is on, I'm signing 9 figure deals on the Uber over here, right? Like that guy isn't going to come and jump on your customer sales calls and like, oh, you know, sure, sure, we'll help you do this. And you know, our team will address this like next week. That's not happening, right? So the level of like customer service that is expected for enterprise SAS is not being provided by the Frontier Labs and they're not in a position to provide it. That's why they started this whole FTE program, but the FDE program, people thought it was a sales engineers program. It's actually a program to extract data and work flows and implement them inside the models themselves. And, and that is kind of, that is what Palantir, what Alex Carp is alluding to. He's like you, you're the FD ES are coming in and they're not there to help you to absorb your work flows. And once you absorb your work flows, you won't have a business because we take it and, and it's true. It's absolutely true. You know, that's what you know, we, we spoke to two opening IFDES and they went in into a company that Thrive owned rather than an external company. And as they went in then they took apart the workflow and they're basically absorbing it. And they said that the intention is to absorb it in the next round. And that is happening right now.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.