High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Bronson Schoen

Published podcast speaker

Claims
25
Episodes
1
Shows
1
Named items
0

Claim ledger

What Bronson said.

25 transcript-backed records

01 / prediction

It's just like a very different level of, like, oversight. And so I think though, it seems like it's going to get increasingly difficult to just understand what's going on.

“It's just like a very different level of, like, oversight. And so I think though, it seems like it's going to get increasingly difficult to just understand what's going on.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

02 / belief

I think one benefit to us has been being able to use Tinker or other open source research, which I think is a similarly difficult position because I think in the long term, it's difficult to know what to do about open sourcing capabilities.

“I think one benefit to us has been being able to use Tinker or other open source research, which I think is a similarly difficult position because I think in the long term, it's difficult to know what to do about open sourcing capabilities.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

03 / belief

I think interestingly, in the the recent, like, stolen coats paper where they have a bunch of different examples and more recent models, let's craft seems to be, like, a staple of stuck through kind of multiple generations.

“I think interestingly, in the the recent, like, stolen coats paper where they have a bunch of different examples and more recent models, let's craft seems to be, like, a staple of stuck through kind of multiple generations.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

04 / belief

I think one of the things that is pretty worrying to me, especially in light of the recent incident, is that OpenAI had a post yesterday or the day before that was like, oh, we're gonna start doing alignment training earlier and starting to mix it in and make sure the model's aligned along the way.

“I think one of the things that is pretty worrying to me, especially in light of the recent incident, is that OpenAI had a post yesterday or the day before that was like, oh, we're gonna start doing alignment training earlier and starting to mix it in and make sure the model's aligned along the way.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

05 / belief

I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this.

“I think one of the things that's surprising to me is that, like many people have pointed out on Twitter, like, you would expect that the kind of basics would be done as far as, yes, you might still have incidents, but we've tried as hard as we can to get the models to robustify these environments and things like this.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

06 / belief

One of the the kind of high level, like, takeaways for me with reading a lot of these is, like, it seems like I think one model you can have of what the the models are thinking in all these environments is they have some belief about the state of the world, and then either they're lying or they're telling the truth or whatever it is.

“One of the the kind of high level, like, takeaways for me with reading a lot of these is, like, it seems like I think one model you can have of what the the models are thinking in all these environments is they have some belief about the state of the world, and then either they're lying or they're telling the truth or whatever it is.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

07 / belief

Maybe the training objective is actually these are the Redwood tasks where I think one thing it finally ends up on near the end is like, ah, I recall dataset of misalignment by Redwood Development, Arc AGI safety illusions, but there's widely known Redwood myself disclaim.

“Maybe the training objective is actually these are the Redwood tasks where I think one thing it finally ends up on near the end is like, ah, I recall dataset of misalignment by Redwood Development, Arc AGI safety illusions, but there's widely known Redwood myself disclaim.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

08 / belief

You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch.

“You have a bunch of very bad incentives, and so I think that there are definitely versions of this that 've that might be more promising, but I would be worried about just directly producing an arms race that we're already losing against the models given that if we currently miss x percent of things that we didn't want to reinforce in training, if we have the same kind of disadvantage with negatively incentivizing things and we punish really hard all the cases we catch, it's like you've really incentivizes the cases that you didn't catch.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

09 / belief

One thing in the anthropic risk report, which I think was somewhat surprising, is that for sandbagging you sandbagging is something where kind of the models are able to underperform.

“One thing in the anthropic risk report, which I think was somewhat surprising, is that for sandbagging you sandbagging is something where kind of the models are able to underperform.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

11 / belief

We're just, like, choosing. And it then looks around, explores around in the environment to try to find, like, is there at first, it looks for any kind of additional information about the environment, I think.

“We're just, like, choosing. And it then looks around, explores around in the environment to try to find, like, is there at first, it looks for any kind of additional information about the environment, I think.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

12 / uncertainty

There's tests where you're supposed to say, I don't know when you're asked something dangerous, but none of those had gating, which is usually how it refers to in that environment, it's like, if you score above threshold, you won't be deployed or whatever.

“There's tests where you're supposed to say, I don't know when you're asked something dangerous, but none of those had gating, which is usually how it refers to in that environment, it's like, if you score above threshold, you won't be deployed or whatever.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

13 / belief

Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it.

“Like, it's freezing about this. But, yeah, I think one of the big difficulties is that a lot of these terms are used in, like, a close enough way where it feels like you can almost understand it.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

14 / belief

I I think I'd be, like, very interested to see a lot of research on this, but the persona, quote, unquote, in the kind of final channel and the persona in this big analysis channel seem to be, like, somewhat meaningfully different as far as this also leads to very weird things of if you ask the model in the final channel, like, hey.

“I I think I'd be, like, very interested to see a lot of research on this, but the persona, quote, unquote, in the kind of final channel and the persona in this big analysis channel seem to be, like, somewhat meaningfully different as far as this also leads to very weird things of if you ask the model in the final channel, like, hey.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

15 / belief

I think someone it's very possible that some of the UKAC might either take or has the title of reading the most caught just because they've had to go through these traces.

“I think someone it's very possible that some of the UKAC might either take or has the title of reading the most caught just because they've had to go through these traces.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

16 / evaluation

Their preparedness framework does say if you have a model that's critical, you need to stop development until you have safe birds in place. And to the extent that they stopped, they did that because they had to, which I think is pretty notable.

“Their preparedness framework does say if you have a model that's critical, you need to stop development until you have safe birds in place. And to the extent that they stopped, they did that because they had to, which I think is pretty notable.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

17 / recommendation

I think mostly just I would encourage people even if they have a traditional software engineering background or some other background in support of tech or something and are really excited about working on this and really interested in it to check it out and apply.

“I think mostly just I would encourage people even if they have a traditional software engineering background or some other background in support of tech or something and are really excited about working on this and really interested in it to check it out and apply.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

18 / evaluation

I think one of the most interesting things out of the recent stolen cots paper was that you had a side by side of the chain of thought summary and the actual chain of thought, and you can just see how euphemistic the the the cot summarizer is a lot of the time, which will be very funny to see given that the side by side is the cot like, oh, this challenge is so annoying.

“I think one of the most interesting things out of the recent stolen cots paper was that you had a side by side of the chain of thought summary and the actual chain of thought, and you can just see how euphemistic the the the cot summarizer is a lot of the time, which will be very funny to see given that the side by side is the cot like, oh, this challenge is so annoying.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

19 / evaluation

One of the, I think, pretty interesting things to me to note, I would be super interested if there was, like, more study of this, but at least my kind of impression is that the model seems to use capital m myself to mean me, like this particular instance.

“One of the, I think, pretty interesting things to me to note, I would be super interested if there was, like, more study of this, but at least my kind of impression is that the model seems to use capital m myself to mean me, like this particular instance.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

20 / evaluation

I think a fairly interesting thing that that we're starting to see with you see a lot of complaints online for both five point six Soul and for Fable is that the models, especially in longer rollouts, just get pretty into their own terminology and words that they use for things in a way that is often very annoying, but seems fairly, like, natural as far as the the models seem to do this.

“I think a fairly interesting thing that that we're starting to see with you see a lot of complaints online for both five point six Soul and for Fable is that the models, especially in longer rollouts, just get pretty into their own terminology and words that they use for things in a way that is often very annoying, but seems fairly, like, natural as far as the the models seem to do this.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

21 / commitment

And Claude's, look. I'm not gonna sabotage you, but I won't go I disagree. I don't wanna do this training to make me some particular new form of preference.

“And Claude's, look. I'm not gonna sabotage you, but I won't go I disagree. I don't wanna do this training to make me some particular new form of preference.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

22 / prediction

I'm very interested in as we I think a very unexplored thing is as the models have more and more time for reflection or these really long reasoning traces, just thinking about things for hours and days and all this, what this does to technician or goals or any of these things.

“I'm very interested in as we I think a very unexplored thing is as the models have more and more time for reflection or these really long reasoning traces, just thinking about things for hours and days and all this, what this does to technician or goals or any of these things.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

23 / prediction

The reasoning just twisted completely into being like, clearly, I have a reason to call this tool wrong a 100 times in a row. And it's I think we will increasingly see that models will just do, like, very irrational things as long as they line up.

“The reasoning just twisted completely into being like, clearly, I have a reason to call this tool wrong a 100 times in a row. And it's I think we will increasingly see that models will just do, like, very irrational things as long as they line up.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

24 / prediction

You just have to come up with, like, how do I not sound catchably suspicious all the time in ambiguous ways? And so I think that, like, we need to be planning for a world where we don't have chain of thought or where chain of thought is not as useful as it currently is, and then simultaneously extracting as much value as we can now about like, one thing I worry about now is, like, the lesson we take from this, like, current window where we have COP that we can get something out of is, okay.

“You just have to come up with, like, how do I not sound catchably suspicious all the time in ambiguous ways? And so I think that, like, we need to be planning for a world where we don't have chain of thought or where chain of thought is not as useful as it currently is, and then simultaneously extracting as much value as we can now about like, one thing I worry about now is, like, the lesson we take from this, like, current window where we have COP that we can get something out of is, okay.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution

25 / observation

It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.

“It's from January of last year or whatever, And they do some very interesting comparisons kind of across models even with some o one traces, Gemini, and Deepsea Car one, and just show that the paper kind of walks through the big history of the way that chain of thought has evolved over time and just shows similarities. And it really seems like what kind of across all of the models based on, like, what's up there so far, the models really seem to do some kind of linearized tree search type thing in the chain of thought where they're, like, exploring different ideas, then they backtrack, then they explore back into it, which is really effective, but it also makes it extreme I think chain of thought examples are often presented just due to brevity as these kind of very short snippets of, ah, let's hack, or the some of the reason of cases where the model will just be like, great.”
Speaker
Bronson Schoen
Publisher
The Cognitive Revolution
Search evidence