Evidence receipt / evaluation
Published · transcript-backedAndreas Stuhlmüller: evaluation
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.”
Source trail
Everything needed to verify it.
- Speaker
- Andreas Stuhlmüller
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Cool. That's really interesting. What do you think is going to drive the Next, if indeed we are successful in having the company continue to function through the holidays without you guys for two weeks, then one wonders like, how long could it go? And do you ever have to come back for one thing? But what will take us there? Is it just the next generation of models? Mythos is supposedly coming soon to a public API near you with supposedly much better long horizon performance. Is that going to be the biggest unlock? You've got the structure, now you just need to drop in a better model, or what else do you think is going to be needed over the next six months to actually realize that? As a first, I don't actually expect the company will like run fully automatically by the end of the year. I expect, so I expect our lower bar is within each function, they're like pretty autonomous workflows that run and that like connect to some workflows in other functions. But I think a lot of the high level steering will still be very much needed. Yeah, what is needed for scale up? I think one big obstacle right now is the models are like not fully calibrated about when human intervention is needed. You have to be like pretty risk averse in how you use them. I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time. That's pretty rough. And so we have to earn much more on the side of, if in doubt, it's not an automatically reviewable feature. And so I think that's the case in software engineering. I expect it's the case in other situations too. Like, you know, if you were to let the models drive through some customer interaction, for example, like you probably want to be at least as sure as in the engineering case. And so it could, it's actually not clear to me how, I think there's like the, if you're following the meter graph, right, there's like the 50% success rate kind of curve that goes up over time as the miles get better, and then there's the 80% success rate curve. And the 50% success rate is like much higher, obviously, than the 80%. And the 80% hasn't been going up quite as fast, I would say, as we would like, and often we want more than 80%. So depending on how the kind of average case performance compares to the kind of, I don't know, 95th percentile performance, just dropping in the next models might be good enough or might not, but I wouldn't automatically rely on it. And it could be, would be cool if there were similar to fast mode, if there were like a ultra reliable mode or something, which isn't just think more, but it's like half guarantees on certain classes of errors that you're never going to make. Yeah, that's cool. A lot of really interesting thinking there. One big question that is generating quite different takes at the moment is, are people going to be able and willing to pay the exponentially rising token bills that the industry as a whole is currently seeing. You could analyze this from any number of ways. One would be like your own internal work, right? Like where is your token budget as compared to your headcount budget today in engineering? And do you expect that With an introduction of mythos, if it really is, let's say, a lot better for some sort of definition of a lot better, do you think you will shift that budget and just spend a lot more on tokens relative to humans compared to what you do today? And then do you think your customers will do that as well? And maybe it will break down by use case. It sure seems like for as much as we do hear a lot of complaining about token costs and anecdotally, oh, this company pulled back or that company hit budget, It still feels to me like there's a lot of value in the marginal intelligence and just getting better results. Tokens are still pretty cheap. In most companies, it's still a small, I hear things like 5, 10% of what we're spending on headcount and that's not that much. Maybe you didn't budget for it and that creates some discomfort in your organization. But on the fundamental economics, it feels to me like if you can just get lots better work, for somewhat even maybe a multiple token cost, still seems like pretty rational to pay it. I guess that's my starting position. What do you guys think you will do? What do you think your customers will do?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.