Evidence receipt / evaluation
Published · transcript-backedJohn Collison: evaluation
24 Mar 2026 Cheeky Pint The 20-year journey to fully autonomous cars with Dmitri Dolgov of Waymo
“LLMs are good at text or tokens, specifically, and obviously perform best at domains that have some single corpus of text they can work on, like coding, where it's very helpful that everything was just textual already.”
Source trail
Everything needed to verify it.
- Speaker
- John Collison
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 Mar 2026
- Publisher
- Cheeky Pint
Transcript context
…Exactly. Then, of course, people like to talk about architectures, but architecture is important, but really a lot of it comes down primarily to your metrics, to your evaluation mechanisms, to all of the training recipes, and of course, data. LLMs are good at text or tokens, specifically, and obviously perform best at domains that have some single corpus of text they can work on, like coding, where it's very helpful that everything was just textual already. Part of the success has been creating textual representations for domains such that we can then put a lens against them. Can you describe how you encode the world that you're seeing? Are you just building a 3D bit map, essentially? This is where I think we get a bit into this question of what is the interface between the encoder and the decoder parts. I think that touches also on the thing you flagged earlier where people like to debate end-to-end or not end-to-end. Let's talk a little bit about end-to-end and then get back to what is the interface between those two. When we say end-to-end, what do we mean? We mean that it is some large ML model. Typically, you don't build them monolithically. You have different parts and different subgroups. But what's important is that you can propagate/back prop the gradient and the loss function all through the different layers. Every layer, you can learn the weights and the representations that matter for the final task. You don't force it through some narrow funnel between, let's say, the encoder and the decoder.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.