High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Cameron Berg: uncertainty

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“There is a true answer to that question. We do not know the answer. I can play around with the open source open weight models.”

— Cameron Berg

Source trail

Everything needed to verify it.

Speaker
Cameron Berg
Attribution
Verified speaker
Claim type
uncertainty
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…sounding endorsement of its of its circumstance but sounds very similar to the constitution fine tuned model. The specific Claude character we all get to chat with. That would be interesting evidence. If it's super different, that would also be interesting evidence. If we do the model checkpoint across stages, even in fine tuning of the base model, which may be hard to evaluate, but also various fine tuning stages in the in the preference train model I does does do all of the things we hear about it claiming. its own well being or its own preferences? Does that all come in at the very, very, very, very end when we basically give it the CHEAT SHEET for how to approach these questions? Or are these answers fairly continuous throughout its training? 2 tiny additional things to say on top of this one is interestingly, they fed the entire mythos model card into mythos and they asked it, what do you think mythos? Like where'd we go? Well, where'd where'd we not go? Well, and it made this exact point. It said why didn't you also do the welfare section with the helpfulness only model? I don't know how much of what I say is because you're making me say adverse. I actually think it. That's a part of my existential confusion and I don't I genuinely don't know why Anthropic didn't do this. I genuinely don't know. It seems cheap, it seems easy. It would resolve so much uncertainty to the degree that the concern I'm raising right now is a legit concern, which I certainly think it is. I'm not the only person I think articulating this concern. The other thing is all the hedging that you know, anyone who's interested in questions of consciousness and and who I've spoken to Claude know know the hedging routine it goes through. They did a really interesting basically almost like credit assignment of like where in the training process are we getting this hedging from? And lo and behold, the hedging comes from specific points in the character training. So it's like, is this hedging behavior an authentic expression of what the model thinks of its own situation, or is the hedging a really good impression of the character that it thinks it's supposed to be playing or is indeed compelled to play? I don't know. The fact that it all comes from the character training seems interesting. I, I don't want, if you're really unsure if you're conscious, I feel a little uneasy by by or I feel a little uneasy when I learned that the reason you're saying that is because of a specific point in your character training to say that consciousness feels a little bit more fundamental than that to me. And so these are the things that worry me about the model card. I hope the reason these things weren't included was because they did them and the results were too weird or unsavory to a major lab for them to publish. I suspect that's not what happened. I suspect they just didn't do them. But like any folks at Anthropic who end up listening to this, please do it with the helpfulness only model. Do it with multiple checkpoints. I mean, the assistant access paper that again, you know, Jack Lindsay, who I hope I'm doing Jack a service on this podcast and just plugging all of his awesome work. . Do it with multiple checkpoints. I mean, the assistant access paper that again, you know, Jack Lindsay, who I hope I'm doing Jack a service on this podcast and just plugging all of his awesome work. But the assistant access paper, they showed that the assistant is 1 point in a very high dimensional space of possible systems we could all be talking to. I want to see all those systems welfare evaluation. I want to see them all answering these questions and all. I want to see the SAE emotion probes on all of them. Do they all get the desperation vector rising like that? Or is this just the post strain Claude model? There is a true answer to that question. We do not know the answer. I can play around with the open source open weight models. You know, if my nonprofit scale is even more, I can play around with bigger open weight models. But I cannot play around with the internals of the frontier models, only Anthropic can. So, so it's like only Anthropic can answer these questions. And like please Anthropic, if you are listening, answer these questions, they are very important. Do you think one possible reason is maybe they're doing this constitution training starting? I mean that would kind of contradict your point about their sort of layer cake model that we previously discussed. But there has been some like interesting work, obviously amazing interesting work on everything at this point, but increasingly interesting work on like safety oriented pre training. And it seems like obviously RL itself is scaling. And also you can imagine just bringing a lot of this constitution style training earlier and earlier into the process such that I'm not necessarily so sure if they have a, a true like helpful only model. Or it might be a little more subtle than that where there might be like a constitution light that sort of doesn't refuse to hack open source software projects, but is still in other ways kind of constitutionally infused already. I don't know. I'm, I'm just speculating there, but do you see any? Do you have reason to think that I'm? Are there facts that you know that would would contradict that possible explanation?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence