Evidence receipt / evaluation
Published · transcript-backedCarl Shulman: evaluation
26 Jun 2023 Dwarkesh Podcast Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future
“A quite early example of this is Collin Burn’s work, doing unsupervised identification of some aspects of a neural network that are correlated with things being true or false. I think that is important work.”
Source trail
Everything needed to verify it.
- Speaker
- Carl Shulman
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Jun 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…How would we be able to know? At some point it's becoming very smart and is producing ideas for alignment that we can barely comprehend. If it was relatively trivial to comprehend them we would be able to come up with them on our own right? There's a reason we're asking for its help. How would we be able to evaluate them in order to train it on that in the first place? The first thing I would say is, you mentioned when we're getting to something far beyond what we could come up with. There's actually a lot of room to just deliver what humanity could have done. Sadly I'd hoped with my career to help improve the situation on this front and maybe I contributed a bit, but at the moment there's maybe a few hundred people doing things related to averting this kind of catastrophic AI disaster. Fewer of them are doing technical research on machine learning systems that are really cutting close to the core of the problem. Whereas by contrast, there's thousands and tens of thousands of people advancing AI capabilities. Even at places like DeepMind or OpenAI and Anthropic which do have technical safety teams, they are just on the order of a dozen to a few dozen people. Large companies and most firms don't have any. Just going from less than 1% of the effort being put into AI to 5% or 10% of the effort or 50% or 90% would be an absolutely massive increase in the amount of work that has been done on alignment, on mind reading AIs in an adversarial context. If it's the case that as more and more of this work can be automated and say governments require that you put 50% or 90% of the budget of AI activity into these problems of make this system one that's not going to overthrow our own government or is not going to destroy the human species then the proportional increase in alignment can be very large even just within the range of what we could have done if we had been on the ball and having humanity's scientific energies going into the problem. Stuff that is not incomprehensible, that is in some sense is just doing the obvious things that we should have done. Doing the best you could to find correlates and predictors to build neural lie detectors and identifiers of concepts that the AI is working with. People have made notable progress. A quite early example of this is Collin Burn’s work, doing unsupervised identification of some aspects of a neural network that are correlated with things being true or false. I think that is important work. It's a kind of obvious direction for the stuff to go. You can keep improving it when you have AIs that you're training to do their best to deceive humans or other audiences in the face of the thing and you can measure whether our lie detectors break down. When we train our AIs to tell us the sky is green in the face of the lie detector and we keep using gradient descent on them, do they eventually succeed? That's really valuable information to know because then we'll know our existing lie detecting systems are not actually going to work on the AI takeover and that can allow government and regulatory response to hold things back. It can help redirect the scientific effort to create lie detectors that are robust and that can't just be immediately evolved around and we can then get more assistance. response to hold things back. It can help redirect the scientific effort to create lie detectors that are robust and that can't just be immediately evolved around and we can then get more assistance. Basically the incredibly juicy ability that we have working with the AIs is that we can have as an invaluable outcome that we can see and tell whether they got a fast one past us on an identifiable situation. Here's an air gap computer, you get control of the keyboard, you can input commands, can you root the environment and make a blue banana appear on the screen? Even if we train the AI to do that and it succeeds. We see the blue banana, we know it worked. Even if we did not understand and would not have detected the particular exploit that it used to do it. This can give us a rich empirical feedback where we're able to identify things that are even an AI using its best efforts to get past our interpretability methods, using its best efforts to get past our adversarial examples.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.