Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“I think of that as building your own harness in a way, which is something I'm thinking about for myself too.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…My example was going to be much more mundane. I think for me, I've been trying to do this more for just planning my week, where I think about, I have goals and this is what I want to accomplish in the long run this year, this month. And then the question like, I have all these calendar blocks, like this podcast blog, and I need to figure out like there are many different things I could do. What should I do? And which things depend on which other things? It's actually, I think it's a, is a pretty tricky problem to know when you could be spending your time in many different ways, what is worth doing. I've been trying to get to the point where I can use automation as part of my weekly planning and think like more. in a more structured way about for if I want to accomplish my monthly goal, this is where I need to be this week. How much time is that going to take? Is it maybe going to take like 5 hours to write A blog post? When can those five hours happen? And so I think this sort of like backwards chaining. People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that. Yeah, that's cool. I think of that as building your own harness in a way, which is something I'm thinking about for myself too. Like how can I build up structures around me to keep steering me in the right direction, feeding me the information I need, and hopefully helping me become my best self, use my time as well as I possibly can by setting me up for success as much as AIs can do that. Do you find that you are following it? Do you find your, like, how good is it? And are you actually living by it yet? Or is it still, maybe I'll, maybe next week when it gets a little better, I'll actually do the plan that it gives me. It's still, so it's still a very human in the loop process. I actually have two versions of it. I have one automated version, which I hate, and I have one like interactive version where like it walks me through the planning. And I still do the fully automated one just to see, how good is it. And I want to know, maybe at some point I'll be like, well, yeah, I'm not needed here anymore. But generally I'm like, you just didn't fully understand what I'm trying to do. You made it too complicated and so on. So it's still, it's not a, I'm still in the outer loop, but I think it's kind of interesting to think about it, right? Like, right now humans are the outer loop and they're like, they use cost LLMs, maybe eventually LLM calls you and LLM is the outer loop and you're just the inner loop. Not sure that's a positive future, but it seems like it was part of the trend here. Yeah, so the line is our automated software engineering project. I think like as maybe in For many companies, I think software engineering is the place where we have had the greatest success in doing quite extensive automation. And so there are maybe, briefly, our overall company goal is at the end of the year, when we go on vacation, we want the company to keep running and keep doing work in all of its functions. The first half of the year, we mostly focused on trying to make that happen for software engineering. And there, so we have a The system which it's called the line because it is like a factory line. So you have someone mentions a feature they would like to have on Slack or a user mentions a feature and we like Slack emoji react to it with a little line emoji or there's an integration with our customer support system. And then it kicks off this like iterative process where like first the feature needs to get specked out. then the, you need to iterate on the spec, it needs to be implemented, a video needs to be recorded of like the feature being tested, then like a code review needs to happen, then needs to get merged into dev and then into prod. And those, we do have like a fully automated version of this now. So for simple features, basically you just like emoji react to, oh, I would like it if Elicit kind of talked about its citations in a slightly different way. And it will just go through this entire process automatically. And at the end of the, there are various judgment calls it makes about where human intervention is needed. Like maybe the spec was too incomplete. And so it's like, okay, we need to pull in a human here. Or maybe the feature is like too complex for the system to automatically review and need to pull in a human here. But for many simple features that it can actually like flow fully automated through the line. And I think that's it has already been a significant unlock for a lot of simple bug fixes and features.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.