Evidence receipt / belief
Published · transcript-backedBronson Schoen: belief
26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
“I think one benefit to us has been being able to use Tinker or other open source research, which I think is a similarly difficult position because I think in the long term, it's difficult to know what to do about open sourcing capabilities.”
Source trail
Everything needed to verify it.
- Speaker
- Bronson Schoen
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 26 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…But Yeah. Well, what do you think of the company's decision decisions to keep chain of thought private at this point? Initially, was like, because they don't wanna expose this ugly stuff that people won't think is pretty because that'll make them Yeah. Pressured to change it. There's also the competitive I think. Insulation. Seems like we could use a lot more eyes on it though too. Right? I think I'm an incredibly bad person to ask because I have the option of being like, it's great that I can see it, but clearly, this should be private from everyone else. But I don't think I think for especially for older models, I think having more eyes on this would be good. And I don't think you can imagine worlds where we're doing some balancing act where we've given a lot of access to people to the cot or whatever it is, and we're trying to figure out, should we get even more access? But right now, it seems like the the number of people with access to this is incredibly small, and it's even a fight to get any kind of third party visibility or auditing into these things. And so I think that getting broader access would be good. I think public access is tricky, I think, but I don't have any good principal reasons for it. I think that the other than just that it does seem to be plausible to me that the model's reasoning, there's good incentives for it not to be as unhinged. I think one of the surprising things to me I don't think it's any my best guess is it's probably not anything nefarious. It's just a bandwidth thing, but I'm surprised at how euphemistic the summarizers are. Like, it doesn't seem like there's a great reason to me that the summarizers couldn't be like, I'm cheating on this task or whatever. It seems like to the extent that you're relying on summarizers, you at least want to trust that the summarizers are to some extent faithful to what is happening. But I'm not sure of a good equilibrium for this. Because I I think there are, like, interventions that you could do from a monitoring perspective that I would be a fan of chain of thought. You just keep doing scratch pad gambits where it's like, you tell the model, hey. You have this persistent workspace that is actually private to only you, and then you see the model write down alignment related things or whatever. But this obviously has the drawback that it's very difficult to ask people to allow models to operate opaque persistent storage on their machines all the time, but I'm not sure of the best solution for these. I think one benefit to us has been being able to use Tinker or other open source research, which I think is a similarly difficult position because I think in the long term, it's difficult to know what to do about open sourcing capabilities. But especially in the short term and especially now, the amount of safety research that gets done on open source models is just incredibly high. Tinker is probably the thing that sped us up the most out of any single thing in the last few years or something. Yeah. I I think it's just a very tricky position to be in. I think one thing that I worry about is the dynamic that you mentioned of more eyes on it of there's both the phenomenon that people can miss things, but also one thing I worry about the labs is that there might just be a lot of things that they're seeing that they just don't have time to investigate. oth the phenomenon that people can miss things, but also one thing I worry about the labs is that there might just be a lot of things that they're seeing that they just don't have time to investigate. If you explain if you explain to a normal person, yeah, our latest model, it says Redwood all the time, but no one's had a chance to look into this. Wait. What the fuck? Why does it always say the one safety or that it knows about? That's really weird. And I've tried to make a similar case to labs about the the multi agent setups because I keep seeing people like, GDM has a good post. It's like, people should study multi agent research more. And then the question is, okay. But what multi agent setup are you guys doing? Because you're only gonna really be interested in results based on the kind of training dynamic setups that you guys have going on. And I think it's it's very feasible for me that capabilities researchers are seeing things that are relevant to alignment on a somewhat regular basis, but just are either kind of frogballed on it or don't really have time to look into it or the alignment team also doesn't have bandwidth. I think one of the stranger things to be seeing with these trillion dollar companies right now is, understandably, you see a lot of the time from safety organizations or safety teams at labs. Oh, yeah. We're just so bandwidth constrained. But there's almost there at a historic level of the number of people who want to work on and be helping with this problem. And so I think it's I'm not sure what the dynamics are, but it seems like there are a lot of things to do for safety and having more people with more eyes on it. Even if the extreme of that has people objecting to it, there are, like, many gradients along the way of allowing more third party access or allowing more people access to older models or whatever it is that seemed really doable to me. And I think that it would be useful if labs thought more in advance of, ah, okay. Whatever is constraining us from hiring right now, it clearly is not money or resources. If we expect to be similarly limited in the future, maybe we should crank up the number of people that we're hiring or or whatever it is. But, yeah, it's very weird to see. It's a very small number of people. I think a very surprising thing to me about this field given the stakes and the money and the the everything involved in it is like when you ask, okay. Who's working on this particular part of alignment that's not going very well? It's, oh, it's four people. It's like, oh, really? It's like five. They got a new person. It's like, okay. Great. But the fact that there's a small number of Rudy Goal that you can just say Ryan, and it's like, oh, yeah. They mean Ryan Greenblatt. There's just mostly a few number of Ryan. I hear this is changing, but it's still surprising to me. But yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.