High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

6 Jun 2026 The Cognitive Revolution AI in the AM — Week 1 Highlights (June 2026)

“A company running thin margins on top of Opus is going to struggle to say, no, don't use that.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…So my take is that we storily need to depend on, well, harmlessness and actual human knowledge is much more important than the models. Cyber Gym, like the most famous cyber evil, the top score right now is by the Microsoft multi-model setup. They used Opus with Sonet and GPT 5.4 and they got a score that is higher than meters. So what we see is that cheaper models can outperform more expensive or smarter models if you just optimize, you know, the knowledge or harness around it. So I think there's a lot of place for humans with real expert knowledge. Let's remember that How to research software is not like a really documented process. It lives in the minds of humans who have been doing this for years. Just like lawyers do, AI agents today, there needs to be somebody sitting there looking at the results and kind of having taste like what is good and what is not. Like at the end of the day, there's somebody behind all of those systems that has to make the judgment call if the quality is up to standard or not. And they will be accountable if something goes wrong, right? You cannot fire an AI somewhere to blame at the end of the day. It's A genuinely good argument. And yet, here's where I come down. When it's security critical, I think people will still pay up for the very best model. A company running thin margins on top of Opus is going to struggle to say, no, don't use that. Use us. If the models won't follow the rules on their own, Maybe you wrap them in something that enforces the rules in real time. Prakash was especially taken with Brett Levinson's pitch for exactly that and with his answer to who the real regulators turn out to be. I would love to dig into the architecture a little bit and then maybe also talk about how this paradigm may extend to things potentially well beyond content policies. On the first point of architecture, I mean, it's got to be fast, right? So like, are you using small models? Is this, is this the sort of thing where you sort of let things through and then run something in the background? And if it gets flagged, then we kind of come in later, like the original Microsoft Bing experience where you'd see the message and then it would like retract it back. Or are you doing the more sort of classifier style approach where it can be fast enough that you can build it into the stack and the latency is acceptable. What trade-offs are people willing to make in terms of product experience, latency, cost, and how are you then engineering to meet their demands? Yeah, so I mean, to me, you've said the magic words. Like I've been a big advocate kind of since we started the company and even since I was at Meta that like an ounce of prevention is worth a pound of cure. Like being there before something happens or as you pointed out, maybe you can optimistically let a message through and then retract it quickly is just a better approach than finding stuff three to seven days later and saying, oh, we screwed up, we need to block or ban this user. And in the case of AI, what would you even do three to seven days later other than maybe like, I don't know, add it as a training example for the next fine tune or something like that? As far as the architecture goes, so we have a couple techniques that we're using. So one, yes, we do use some very small models that are already pretty fast. It also turns out that breaking a policy down in the way we do into atomized bytes gives us some unique advantages on sort of the latency front. The questions we're asking are all pretty small. They tend to share a prefix, basically. And so we're able to sort of benefit from quite a large amount of prefix caching. We also, generally speaking, at least first pass, we're not generating much. There's really no decode step for us. What we, I mean, I'm happy to share some of the architectural details. Like we essentially are training a binary classification head onto an LLM, right? We don't initially anyway need some, we don't need the questions answered with an actual yes or no. And in fact, it's counter to our objectives. to do so, we actually want to know what is the probability that the answer to this question is yes, basically. And that I don't want to get, I don't want to, I have a tendency to sometimes go on tangents. So I'm going to try to contain myself here and maybe we can come back to like the benefits of having those probabilities and the abstained gap and all that. There's another common thing in moderation safety guardrails control, whatever you want to call it, which is that for most, for the majority of policies, upwards of 90% of all the content you're ever going to see is fine. Like it's a real needle in a haystack problem, right? Like you're looking for a small sliver. The only problem is that very often that sliver has high severity, has real risk associated with it. And so we have a number of layers sort of in front, you mentioned lightweight classifiers. They're not simple binary classifiers, but we do have a number of much lighter weight models that sit in front of our, I guess what I would call like our main QA engine that can give us with reasonable confidence and high recall, that's the important part, a quick answer up front. And so the idea is like, For, let's say, just for argument's sake, let's say 90% of what we're gonna get sent from a particular customer is fine. Really, there's no problem. We don't need to look at it for real, basically.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence