High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Adam Gleave

Published podcast speaker

Claims
14
Episodes
1
Shows
1
Named items
0

Claim ledger

What Adam said.

7 transcript-backed records

01 / evaluation

Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.

“Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

02 / evaluation

I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.

“I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

03 / evaluation

When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.

“When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

04 / evaluation

The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.

“The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

05 / evaluation

You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.

“You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

06 / evaluation

Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.

“Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

07 / evaluation

Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input.

“Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution
Search evidence