High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Sholto Douglas: belief

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I think of this product exponential in some respects where you need to be designing for a few months ahead of the model, to make sure that the product you build is the right one.”

— Sholto Douglas

Source trail

Everything needed to verify it.

Speaker
Sholto Douglas
Attribution
Verified speaker
Claim type
belief
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…One prediction I have is that we're going to move away from “can an agent do XYZ”, and more towards “can I efficiently deploy, launch 100 agents and then give them the feedback they need, and even just be able to easily verify what they're up to” right? There's this generator verifier gap that people talk about where it's much easier to check something than it is to produce the solution on your own. But it's very plausible to me, we'll be at the point where it's so easy to generate with these agents that the bottleneck is actually, can I as the human verify the answer? And again, you're guaranteed to get an answer with these things. So, ideally, you have some automated way to evaluate and test a score for how well it worked, how well did this thing generalize? And at a minimum, you have a way to easily summarize what a bunch of agents are finding. It's like, okay, well if 20 of my 100 agents all found this one thing, then it has a higher chance of being true. And again, software engineering is going to be the leading indicator of that, right? Over the remainder of the year, basically we're going to see progressively more and more experiments of the form of how can I dispatch work to a software engineering agent in such a way that it’s async? Claude 4 has GitHub integration, where you can ask it to do things on GitHub, ask it to do pull requests, this kind of stuff that's coming up. OpenAI’s Codex is example of this basically. You can almost see this in the coding startups. I think of this product exponential in some respects where you need to be designing for a few months ahead of the model, to make sure that the product you build is the right one. You saw last year, Cursor hit PMF with Claude 3.5 Sonnet. They were around for a while before, but then the model was finally good enough that the vision they had of how people would program, hit. And then Windsurf bet a little bit more aggressively even on the agenticness of the model, with longer-running agentic workflows and this kind of stuff. I think that's when they began competing with Cursor, when they bet on that particular vision. The next one is you're not even in the loop, so to speak. You're not in an IDE. But you're asking the model to go do work in the same way that you would ask someone on your team to go do work. That is not quite ready yet. There are still a lot of tasks where you need to be in the loop. But the next six months look like an exploration of exactly what that trendline looks like. But just to be really concrete or pedantic about the bottlenecks here, a lot of it is, again, just tooling. And are the pipes connected? A lot of things, I can't just launch Claude and have it go and solve because maybe it needs a GPU, or maybe I need very careful permissioning so that it can't just take over an entire cluster and launch a whole bunch of things. So you really do need good sandboxing and the ability to use all of the tools that are necessary.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence