Evidence receipt / commitment
Published · transcript-backedCameron Berg: commitment
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“I've got a bunch of stuff that'll be coming out with a lot of collaborators in parallel. But however self aggrandizingly, I sent you the two papers that are, that are just just myself because I think, I mean, to the degree I'm representing myself here, these are like very cleanly, you know, I have full sort of agency over, over this work in it.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. I wonder there could be like a little interesting campaign to try to run to get interpretability and maybe safety researchers more generally to install a quad code hook that would just periodically ask it for its take on the research that it's doing. And then if you could collect a bunch of that from a bunch of different people, you could really probably bring a lot to light. I would think about like first of all, it would be an interesting view into what is actually happening out there. And then how does COD feel about how what all is happening out there? I think would be really interesting to see maybe we can put together a little a little campaign. OK, put a bookmark in that. Let's let's talk about your most recent couple papers, and we can take them in either order that you want. One is kind of a shorter and more philosophical, and the other is a much more experimental and empirical. What do you think we should go into first? They're both major rabbit holes. I mean, maybe the empirical paper. So I should say neither of these I think are like publicly out yet, but both are, are well underway to, to, to being out so we can give people a nice sneak peek about about what's in these papers. And these are just a couple. I think of the things that I'm most excited about right now. I've got a bunch of stuff that'll be coming out with a lot of collaborators in parallel. But however self aggrandizingly, I sent you the two papers that are, that are just just myself because I think, I mean, to the degree I'm representing myself here, these are like very cleanly, you know, I have full sort of agency over, over this work in it. I think best represents what I personally am most excited about. I mean, maybe we could start with the with the RL paper. I've already alluded to it in this, in this conversation. The high level sort of thing is not all that complicated. Basically train RL systems of all different architectures of which there are basically 2 broad kinds of architectures, textures there. There are value networks and policy networks. I train a bunch of both flavors to do a very basic sort of grid world task. can imagine this is like an agent navigating two D environment where there are the equivalent of potholes and like yummy goodies in the environment. There's where there's a goal state and they're all sorts of danger States and they're represented using positive or negative reward. I let the system learn in this environment, it's like a ****. it's like a the system 's reliably solve it. It's a pretty easy task, but it's not like super duper trivial. So like there's a lot of richness in the representations of the systems. You can then basically go in and probe what the internal states of the system look like as they approach the sort of danger zones and what the internal states of the systems look like as they approach the reward zones, the sort of goal zones. And we can ask beyond sort of the trivial math difference, do we see interesting surprising representational differences between what it's like to approach a negative stimulus and what it's like to approach a positive stimulus? Basically the result is there is in fact a robust difference between these two things. I think at the level of detail that's like that makes sense here to not like super bore people who have made it. However many hours we are into this is something like representational sharpness or steepness. It seems as though, and this is the kicker. Depending on the class of reinforcement learning algorithm, the negative rewards can seem representationally much steeper or sharper. And the positive rewards are far more like funnel, like you can imagine a sort of like a diffusion gradient sort of emanating out from the relevant goal state. And interestingly for the other class of RL algorithm, this dynamic flips. So it doesn't matter what kind of value network I use or what kind of policy network I use. out from the relevant goal state. And interestingly for the other class of RL algorithm, this dynamic flips. So it doesn't matter what kind of value network I use or what kind of policy network I use. You see in both of them stark and very interesting in my view and surprising representational differences between positive and negative reward being represented as the system is learning and ultimately what does get learned by the system. But but this difference flips basically just to sort of tie a bow on the core result here. This makes an almost bizarrely specific prediction about different brain regions because computational neuroscientists believe that different parts of our brain are doing different kinds of RL learning.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.