High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Boris Cherny: belief

19 Feb 2026 Lenny's Podcast Head of Claude Code: What happens after coding is solved | Boris Cherny

“We released Claude Code really early, because we wanted to study safety. And we actually used it within Anthropic for I think four or five months, or something before we released it, because we weren't really sure.”

— Boris Cherny

Source trail

Everything needed to verify it.

Speaker
Boris Cherny
Attribution
Verified speaker
Claim type
belief
Recorded
19 Feb 2026
Publisher
Lenny's Podcast

Transcript context

…Yeah. I think that point is so interesting, and it's so unique. There's always been this idea, release early, learn from users, get feedback, iterate. The fact that it's hard to even know what the AI is capable of, and how people will try to use it is a unique reason to start releasing things early. So, that'll help you as you exactly describe this idea of what is a latent demand in this thing that we didn't really know? Let's put it out there and see what people do with it. Yeah. And for Anthropic as a safety lab, the other dimension of that is safety, because when you think about model safety there's a bunch of different ways to study it. The lowest level is alignment, and mechanistic interpretability. So, this is when we train the model we want to make sure that it's safe. We, at this point, have pretty sophisticated technology to understand what's happening in the neurons to trace it. And so, for example, if there's a neuron related to deception, we're starting to get to the point where we can monitor it and understand that it's activating. And so, this is alignment, this is mechanistic interpretability, it's the lowest layer. The second layer is evals, and this is, essentially, a laboratory setting, the model is in a Petri dish, and you study it. And you put in the synthetic situation and just say, "Okay. Model, what do you do?" And, "Are you doing the right thing? Is it aligned? Is it safe?" And then the third layer is seeing how the model behaves in the wild. And as the model gets more sophisticated, this becomes so important, because it might look very good on these first two layers, but not great on the third one. We released Claude Code really early, because we wanted to study safety. And we actually used it within Anthropic for I think four or five months, or something before we released it, because we weren't really sure. Like, this is the first big agent that I think folks had released at that point. It was definitely the first coding agent that became broadly used. And so, we weren't sure if it was safe. And so, we actually had to study it internally for a long time before we felt good about that. And even since there's a lot that we've learned about alignment. There's a lot that we've learned about safety, that we've been able to put back into the model, back into the product. And for Cowork, it's pretty similar. The model is in this new setting. It's doing these tasks that are not engineering tasks. It's an agent that's acting on your behalf. It looks good on alignment, it looks good on evals, we tried it internally, it looks good. We tried it with a few customers, it looks good. Now we have to make sure it's safe in the real world. And so, that's why we release a little early. That's why we call it a research preview. But, yeah. It's constantly improving. And this is really the only way to make sure that over the long-term the model is aligned, and it's doing the right things. It's such a wild space that you work in where there's this insane competition and pace. At the same time, there's this fear that if the God can escape and cause damage, and just finding that balance must be so challenging. What I'm hearing is there's these three layers, and I know there's ... This could be a whole podcast conversation is how you all think about the safety piece, but just what I'm hearing is there's these three layers you work with. There's observing the model thinking and operating. There's tests, evals that tell you this is doing bad things. And then releasing it early. I haven't actually heard a ton about that first piece. That is so cool. So, you guys can ... There's an observability tool that can let you peak inside the model's brain and see how it's thinking and where it's heading.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence