High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Lambert: evaluation

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT.”

— Nathan Lambert

Source trail

Everything needed to verify it.

Speaker
Nathan Lambert
Attribution
Verified speaker
Claim type
evaluation
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…This is why a lot of models today, even if they train on zero OpenAI data, you ask the model, “Who trained you?” It’ll say, “I’m ChatGPT trained by OpenAI,” because there’s so much copy paste of OpenAI outputs from that on the internet that you just weren’t able to filter it out and there was nothing in the RL where they implemented or post-training or SFT, whatever, that says, “Hey, I’m actually a model by Allen Institute instead of OpenAI.” We have to do this if we serve a demo. We do research and we use OpenAI APIs because it’s useful and we want to understand post-training and our research models, they all say they’re written by OpenAI unless we put in the system prop that we talked about that, “I am Tülu. I am a language model trained by the Allen Institute for AI.” And if you ask more people around industry, especially with post-training, it’s a very doable task to make the model say who it is or to suppress the OpenAI thing. So in some levels, it might be that DeepSeek didn’t care that it was saying that it was by OpenAI. If you’re going to upload model weights, it doesn’t really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to tailor it to that and it doesn’t matter that it’s saying it’s ChatGPT. Oh, I guess one of the ways to do that is like a system prompt or something like that? If you’re serving it to say that you’re-…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence