Evidence receipt / commitment
Published · transcript-backedAlexander Panfilov: commitment
22 Aug 2026 Machine Learning Street Talk Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov
“I mean, I think also opening eyes have this after all this into Zen's, like, now we are expanding our, like, train of thought monitors and, like, we're putting more effort into it. And, yeah, I think we need just, you know, do more safety mitigations, do more monitoring, see what's model is up to, try to see where it's come from, and maybe we can mitigate it.”
Source trail
Everything needed to verify it.
- Speaker
- Alexander Panfilov
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 22 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. But what was the question? Sorry. Well, you know, Ilya Ilya was just saying that, you know, we we shouldn't anthropomorphize. You know, like, I was I interviewed the Poly Research a couple of weeks ago, and, you know, they were talking about this phenomenon of of reward seeking. Yes. And and they said it's distinct from reward hacking because the model could can conceptualize the reward environment, which is super interesting. Right? Because, you know, they're they're reinforced with these RL traces. So it doesn't explicitly know about the the, you know, like, concept of a grader, but it learns to conceptualize it. Yeah. And and they're saying that the models are sort of, like, you know, becoming agentic and and sort of, like, learning these very abstract concepts in a in a similar way to how we do. And the evidence seems to support it at least in some way. I mean, I think it's it's a definitely frontier research what Apollo is doing and that good that they're looking into it. And I think we I mean, I think also opening eyes have this after all this into Zen's, like, now we are expanding our, like, train of thought monitors and, like, we're putting more effort into it. And, yeah, I think we need just, you know, do more safety mitigations, do more monitoring, see what's model is up to, try to see where it's come from, and maybe we can mitigate it. Yeah. But I think I agree with you, Leo, on this. Like, it would be nice to have some controlled environments and, like, maybe some have counterfactuals like that. If we haven't done this in our training pipeline, what if this happened? Or, like, if model was not the well aware or if it was a well aware and, like, how it contributes to the thing. So Yeah. I think it's just we're a bit too poor compute wise. If we could properly study this, and maybe eventually we'll get to a point where we can. It's but but it definitely requires a very precise experiment. Like, as a scientist, it just feels very hard to say, no. No. No. This is exactly this is the phenomenon. That's it. No. It's very observational studies. You can't prove a hypothesis. You can only reject hypothesis. Right? Like, it's the very fundamental truth of all of this. We are just observers. So let's see. Let's see what happens. Well, apparently, Nathan Lambert said,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.