High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Jérémy Scheurer

Published podcast speaker

Claims
5
Episodes
1
Shows
1
Named items
0

Claim ledger

What Jérémy said.

5 transcript-backed records

01 / belief

You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.

“You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.”
Publisher
Machine Learning Street Talk

02 / belief

Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.

“Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.”
Publisher
Machine Learning Street Talk

03 / prediction

We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.

“We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.”
Publisher
Machine Learning Street Talk

04 / evaluation

For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning.

“For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning.”
Publisher
Machine Learning Street Talk

05 / observation

This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.

“This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.”
Publisher
Machine Learning Street Talk
Search evidence