Evidence receipt / evaluation
Published · transcript-backedRyan Lopopolo: evaluation
7 Apr 2026 Latent Space Extreme Harness Engineering for Token Billionaires: 1M LOC, 1B toks/day, 0% human code, 0% human review — Ryan Lopopolo, OpenAI Frontier & Symphony
“Yeah, and I think success for the way we are [01:08:00] operating here means the model gets better taste because we can point the way there and none of the things we have built actively degrade Asian performance.”
Source trail
Everything needed to verify it.
- Speaker
- Ryan Lopopolo
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 7 Apr 2026
- Publisher
- Latent Space
Transcript context
…you have this too, you deploy a billion tokens of intelligence a day and this was, this was at the beginning of 2096. You’re Yeah. Cooking. Yeah, there’s this fundamental tension, which I think you have talked about between whether or not we invest deeper into the harness or we invest deeper into the training process to get the model to do more of this by default. Yeah, and I think success for the way we are [01:08:00] operating here means the model gets better taste because we can point the way there and none of the things we have built actively degrade Asian performance. ‘cause really all they’re doing is running tests and like running tests is a good part of what it means to write reliable software. If we were building an entire separate rust scaffold around Codex to restrict its output, that I think would be like additional harness that would be prone to being scrapped. But yeah. Yeah. If instead we can build all the guardrails in a way that’s just native to the output that Codex is already producing, which is code, I think. No friction with how the model continues to advance, but also like just good engineering and that’s the whole point. Yeah. So I’ve had similar discussions with research scientists where the RL equivalent is on policy versus off policy. Yeah. And you’re basically saying that you should build an on policy harness, which is already within distribution and you modify from there. But if you build it off policy, it’s not that useful.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.