Evidence receipt / recommendation
Published · transcript-backedCameron Berg: recommendation
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“Keenan actually deserves a big shout out here too because he has another paper called Selfie, which I won't get into the details, but it basically allows you to bootstrap SAE labels so that you can just have way more accurate labels on your SAE given basically having the model label its own activations.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Is that powered again by the good Fire API, same as you had used last time? Yeah, exactly. And the good folks at Studio actually built a replacement to the Good Fire API because the folks at Good Fire retired their API somewhat abruptly. As much as I love the work they're doing, I was. I and other mech and terp flavored researchers were pretty sad to see them. Sort of just make the API disappear. So while I was still doing my work at AEI and a couple other people were sort of very motivated to to basically rebuild the good fire API. we took the same llama seventy B SAE that they trained and found a way to serve it via API which everyone can actually go to it's steering API dot com and i think anyone can go use it. And so that's that's what they might have used good fire when they did this work. But if you want to do it or replicate it, or for that matter, you know, replicate my deception paper or anything like that, you can basically use the same API. Keenan actually deserves a big shout out here too because he has another paper called Selfie, which I won't get into the details, but it basically allows you to bootstrap SAE labels so that you can just have way more accurate labels on your SAE given basically having the model label its own activations. It is also a little introspection flavored, but you basically can end up with better labels. And you started with an SAE by having a model label, the nature of what you're activating by basically feeding it a soft token rather than feeding it a language. You can be like, you know, the capital of of France is this sort of vector, the soft token, and then it will be able to sort of label that itself. And so anyway, this we use the selfie labels on steering API. So the labels are even better than than what good fire offered. So anyway, that's the tooling that that we're using it and the tooling I continue to use. I think it's an excellent, excellent tool for people to play around with. Yeah, cool. well i mean llama seven DB is not you know it's pretty llama three seventy B right. It's pretty far from the frontier. So it is striking to see that the these things are happening already at that scale. I guess a couple things I'd like to try to get a better understanding of at least your intuition for if, if there's not anything that we could consider like a canonical or fully evidence based understanding. 1 is how do we connect these abilities to the idea that there is an experience of these abilities. I mean, it's a striking ability that that models can do this. It's surprising in the sense that I highly doubt this was ever trained for to create review see any evidence to the contrary. But I would my strong assumption would be that Llama three training did not include any incentive and any any reward or any, you know, any gradient descent pushing it toward. we have seen by the way in other papers like activation oracle 's that you can train models to do this also pretty readily as well. That's maybe a little less shocking and and in some ways like potentially really useful. But this is seemingly something that is happening spontaneously, not because anybody intended for it to happen. And I guess, yeah, maybe. So maybe 2 questions are like how do we understand why this would be happening at all? You know, it's, it seems quite surprising, but even now that we've seen it, do we have a theory and we've got this additional detail of it seems to happen more or only under certain preference based tuning as opposed to purely imitative learning? Do we have a story that we find compelling as to why one training paradigm would give rise to these features while the other one doesn't? And then on top of that, like, how do we think about the relationship between this and actual experience? How would you respond to somebody who says, like, that's amazing that that happens, and I'm surprised to see it. But I still don't share your intuition that this has much bearing on whether I should think the, you know, models are ultimately experiencing something that I should care about, you know, in sort of a moral, patient sense.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.