High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

John Collison: evaluation

24 Mar 2026 Cheeky Pint The 20-year journey to fully autonomous cars with Dmitri Dolgov of Waymo

“I think of a simple view of end-to-end being pixels go in, and car actions come out, which may be a bit of an oversimplification.”

— John Collison

Source trail

Everything needed to verify it.

Speaker
John Collison
Attribution
Verified speaker
Claim type
evaluation
Recorded
24 Mar 2026
Publisher
Cheeky Pint

Transcript context

…This is where I think we get a bit into this question of what is the interface between the encoder and the decoder parts. I think that touches also on the thing you flagged earlier where people like to debate end-to-end or not end-to-end. Let's talk a little bit about end-to-end and then get back to what is the interface between those two. When we say end-to-end, what do we mean? We mean that it is some large ML model. Typically, you don't build them monolithically. You have different parts and different subgroups. But what's important is that you can propagate/back prop the gradient and the loss function all through the different layers. Every layer, you can learn the weights and the representations that matter for the final task. You don't force it through some narrow funnel between, let's say, the encoder and the decoder. I think of a simple view of end-to-end being pixels go in, and car actions come out, which may be a bit of an oversimplification. That's exactly right. This is the basic vanilla version of it. If you think about what will it take to build the driver that's capable of fully autonomous operations. You think about this entire ecosystem of the driver, the simulator, the critic. If that's all you do—pixels in, trajectories out—it becomes very difficult to do all of those three and achieve the high level of safety and performance that we require, and it becomes very difficult to do it at scale. However, it's a very easy way to get started. You collect some data… Kind of like the LLM world. The easiest thing you can do is pick a model. The easiest way to get started nowadays would be just take a VLM. It already has a language-aligned camera encoder, and then it has a decoder that can predict, generate text, and you can fine-tune it and say, "Instead of text, generate trajectories." Very doable. In fact, a while ago, we published a paper called EMMA, that did exactly that. It will actually, in the nominal case, drive pretty darn well, which is mind-blowingly impressive.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence