High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“I think would be really interesting to see maybe we can put together a little a little campaign.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…thing that I have. I was doing a similar work along these lines with a couple people. And you know, Owen and Yan definitely scooped us and did a way better version of what we were playing around with. But I saw similar things on my end and playing, playing with the same experiment of basically get the model to believe it's conscious and then see what else comes along for the ride and all sorts of very interesting and obvious things. Interesting. Some obvious, some less obvious things come along for the ride. And yeah, I agree, that's a really interesting area of research that that we should all be paying more attention to. Because again, the direction does seem to just be going in one way here. The the credences and model consciousness seem to be monotonically increasing. And So what happens when we enter a world where, you know, either the models themselves believe them, believe themselves to be conscious, or lots of people or the relevant kinds of people believe the models are conscious or some combination of those two things. What does that world look like? It's an incredibly interesting question. I don't have the answer to it. i think it's it's a lot 's going to change pretty quickly. And I really what I do feel confident about is that us being proactive and thinking through these things will make that world go better than if we basically just sweep the thing under the rug. We can get away with doing it because we still have full control over how all this is going while simultaneously passing off, as you allude to a lot of major decisions and how we're building these systems to the systems themselves. That is only going to keep happening happening with recursive self improvement. As you're saying, it's already happening. I know folks at the major labs are using the best versions of their current models to help build the next versions of the models. you know the trivial example is that some hundred percent of clawed code was written using clawed code according to the guy who's who's who's leading on clawed code. And so this is already happening. And yeah, I think we just probably it would be wise to be proactive about this rather than wait for the models to to be in control of these decisions. And then they're like, well, when humanity was in control, no one really thought carefully about this. So we'll take it from here. Thanks a whole lot, guys. I don't want to be in that world. And so maybe this is just like a long winded way of dodging your question. But at the very least I don't have a good answer for you right now. I don't think anybody does. And I think we better start thinking about it pretty damn soon if we want the long term future to go well with these systems. Yeah. I wonder there could be like a little interesting campaign to try to run to get interpretability and maybe safety researchers more generally to install a quad code hook that would just periodically ask it for its take on the research that it's doing. And then if you could collect a bunch of that from a bunch of different people, you could really probably bring a lot to light. I would think about like first of all, it would be an interesting view into what is actually happening out there. And then how does COD feel about how what all is happening out there? I think would be really interesting to see maybe we can put together a little a little campaign. OK, put a bookmark in that. Let's let's talk about your most recent couple papers, and we can take them in either order that you want. One is kind of a shorter and more philosophical, and the other is a much more experimental and empirical. What do you think we should go into first? They're both major rabbit holes. I mean, maybe the empirical paper. So I should say neither of these I think are like publicly out yet, but both are, are well underway to, to, to being out so we can give people a nice sneak peek about about what's in these papers. And these are just a couple. I think of the things that I'm most excited about right now. I've got a bunch of stuff that'll be coming out with a lot of collaborators in parallel. But however self aggrandizingly, I sent you the two papers that are, that are just just myself because I think, I mean, to the degree I'm representing myself here, these are like very cleanly, you know, I have full sort of agency over, over this work in it. I think best represents what I personally am most excited about. I mean, maybe we could start with the with the RL paper. I've already alluded to it in this, in this conversation. The high level sort of thing is not all that complicated. Basically train RL systems of all different architectures of which there are basically 2 broad kinds of architectures, textures there. There are value networks and policy networks. I train a bunch of both flavors to do a very basic sort of grid world task. can imagine this is like an agent navigating two D environment where there are the equivalent of potholes and like yummy goodies in the environment. There's where there's a goal state and they're all sorts of danger States and they're represented using positive or negative reward. I let the system learn in this environment, it's like a ****. it's like a the system 's reliably solve it. It's a pretty easy task, but it's not like super duper trivial. So like there's a lot of richness in the representations of the systems. You can then basically go in and probe what the internal states of the system look like as they approach the sort of danger zones and what the internal states of the systems look like as they approach the reward zones, the sort of goal zones. And we can ask beyond sort of the trivial math difference, do we see interesting surprising representational differences between what it's like to approach a negative stimulus and what it's like to approach a positive stimulus? Basically the result is there is in fact a robust difference between these two things. I think at the level of detail that's like that makes sense here to not like super bore people who have made it. However many hours we are into this is something like representational sharpness or steepness. It seems as though, and this is the kicker. Depending on the class of reinforcement learning algorithm, the negative rewards can seem representationally much steeper or sharper. And the positive rewards are far more like funnel, like you can imagine a sort of like a diffusion gradient sort of emanating out from the relevant goal state. And interestingly for the other class of RL algorithm, this dynamic flips. So it doesn't matter what kind of value network I use or what kind of policy network I use.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence