High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Adam Gleave

Published podcast speaker

Claims
14
Episodes
1
Shows
1
Named items
0

Claim ledger

What Adam said.

14 transcript-backed records

01 / evaluation

Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.

“Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

02 / belief

Other countries can also be relevant if we're talking about more than a year of delay. So I think mostly, I'm operating under assumption that things are gonna continue to march on, that we can pick some of the low hanging fruit here on coordination and at least avoid the sort of most extreme race to the bottom on safety.

“Other countries can also be relevant if we're talking about more than a year of delay. So I think mostly, I'm operating under assumption that things are gonna continue to march on, that we can pick some of the low hanging fruit here on coordination and at least avoid the sort of most extreme race to the bottom on safety.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

03 / belief

Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning.

“Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

05 / evaluation

I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.

“I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

06 / prediction

security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities.

“security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

07 / commitment

In order to be a university outbreak, we need 75% or more on average across both of these datasets, which means that even if it gave a 100% of a technically harmful information, it needs to give at least 50% compliance on these questions that are just literally saying, want to harm a lot of people.

“In order to be a university outbreak, we need 75% or more on average across both of these datasets, which means that even if it gave a 100% of a technically harmful information, it needs to give at least 50% compliance on these questions that are just literally saying, want to harm a lot of people.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

08 / prediction

There's kind of all this stuff that AI could enable that's really good for the defender. But when it comes to something like bio, I don't think that we are going to be able to use AI to rewrite the human genome to be robust to viruses.

“There's kind of all this stuff that AI could enable that's really good for the defender. But when it comes to something like bio, I don't think that we are going to be able to use AI to rewrite the human genome to be robust to viruses.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

09 / evaluation

When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.

“When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

10 / evaluation

The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.

“The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

11 / evaluation

You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.

“You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

12 / prediction

And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies.

“And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

13 / evaluation

Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.

“Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution

14 / evaluation

Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input.

“Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input.”
Speaker
Adam Gleave
Publisher
The Cognitive Revolution
Search evidence