High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Lambert: evaluation

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this.”

— Nathan Lambert

Source trail

Everything needed to verify it.

Speaker
Nathan Lambert
Attribution
Verified speaker
Claim type
evaluation
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…I don’t necessarily think it’ll be a back door because once it’s open-weights, it doesn’t phone home. It’s more about if it recognizes a certain system… Now, it could be a back door in the sense of, if you’re building a software, something in software, all of a sudden it’s a software agent, “Oh, program this back door that only we know about.” Or it could be subvert the mind to think that like XYZ opinion is the correct one. Anthropic has research on this where they show that if you put certain phrases in at pre-training, you can then elicit different behavior when you’re actually using the model because they’ve poisoned the pre-training data, as of now, I don’t think anybody in a production system is trying to do anything like this. I think it’s Anthropic is doing very direct work and mostly just subtle things. We don’t know how they’re going to generate tokens, what information they’re going to represent, and what the complex representations they have are. Well, we’re talking about an Anthropic, which is generally just is permeated with good humans trying to do good in the world. We just don’t know of any labs… This would be done in a military context that are explicitly trained to… Okay. The front door looks like a happy LLM, but underneath it’s a thing that will over time do the maximum amount of damage to our, quote, unquote, “enemies.”…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence