High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Axel Højmark

Published podcast speaker

Claims
8
Episodes
1
Shows
1
Named items
0

Claim ledger

What Axel said.

7 transcript-backed records

01 / belief

I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.

“I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

02 / belief

What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.

“What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

03 / belief

When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.

“When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

04 / belief

Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.

“Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

06 / belief

Reward hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. So, like, a classic example is, like, in CoastRunners, I think the game is, where the the boat is optimized or, like, the RL agent is optimized to drive a boat in certain in in a in a race, and essentially, what it learns is just to drift in a corner over and over and maximize reward that way.

“Reward hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. So, like, a classic example is, like, in CoastRunners, I think the game is, where the the boat is optimized or, like, the RL agent is optimized to drive a boat in certain in in a in a race, and essentially, what it learns is just to drift in a corner over and over and maximize reward that way.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

07 / belief

Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.

“Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk
Search evidence