evaluation · 13 Feb 2026 · 1:22:06

The claim that continual learning is a significant challenge is often overstated, and the focus should shift to foundational capabilities like code generation.

I feel like that is AGI-complete, which maybe is internally consistent. But it's not like saying 90% of code or 100% of code. No, I gave this spectrum: 90% of code, 100% of code, 90% of end-to-end SWE, 100% of end-to-end SWE. New tasks are created for SWEs. Eventually those get done as well. It's a long spectrum there, but we're traversing the spectrum very quickly.

Watch at 1:22:06