High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Llion Jones: belief

23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]

“Like, this is their job. And what was perfect, I realized, is they tell you in agonizing detail exactly what reasoning they're used to solve those particular puzzles.”

— Llion Jones

Source trail

Everything needed to verify it.

Speaker
Llion Jones
Attribution
Verified speaker
Claim type
belief
Recorded
23 Nov 2025
Publisher
Machine Learning Street Talk

Transcript context

…Right. Yes. So I I wanted to tell you a little bit about this benchmark, because I think I've been having a little bit of issue promoting it, because it doesn't, on the surface, sound particularly interesting. Because sudoku has a sort of a feeling that it's already been solved. Right? So how interesting can a collection of of sudokus be for reasoning exactly? We're not talking about normal sudokus. We're talking about variant sudokus. And what variant sudokus are are usually normal sudokus. Right? So put the numbers 1 to 9 in the row of the column in the box. But then literally any additional rules on top of that. And they're all handcrafted. They all have extremely different constraints. Constraints that actually require very strong natural language understanding. So for example, there's 1 puzzle in the datasets where it tells you the constraints of the puzzle in natural language, and then says, oh, by the way, 1 of the numbers in that description is wrong. Right? So you have to be able to meta reason about the rules themselves even before you start solving the puzzle. There are other puzzles where you have a a maze overlaid on the sudoku, and the rat has to work out a way through the maze by following a path to the cheese, but then there are constraints on the path that it takes of, like, what numbers and what they can be add up to. It's difficult to really describe how varied these these these variants, Sudoku's are. And I think they're so varied that if anyone was actually be able to beat our benchmark, they would necessarily have to have created an extremely powerful reasoning system. Right now, the best models get around 15%, but they're only the very, very simplest and the very smallest Sudoku puzzles in in the set. We're gonna be putting out a blog post about GPT five's performance, and it is a jump, but it's still completely unable to solve puzzles which are, you know, humans can can solve. And what I really like about this data datasets, and actually was the catalyst for me creating it in the first place, it was that there was a there was a quote by Andre Kapathi saying, okay. So we have all this data. It's from the Internet. But what you really want right? If you wanted AGI, you wouldn't want all of the text that humans have ever created. You would actually want the thought traces in their head as they were creating the text. Right? If you could actually learn from that, then you would get something really powerful. And I thought to myself, well, that data must exist somewhere. My first thought was maybe philosophy? Like, you know, there's there's a type of philosophy where you just write down your thoughts without thinking, like just stream of consciousness. I thought, maybe that could work. But then when I wasn't thinking about it, and I was, you know, in my leisure time, I was watching a YouTube channel called cracking the cryptic Yes. of consciousness. I thought, maybe that could work. But then when I wasn't thinking about it, and I was, you know, in my leisure time, I was watching a YouTube channel called cracking the cryptic Yes. Where these these 2 British gentlemen will solve these extremely difficult Sudoku puzzles for you. Like, sometimes their their videos are 4 hours long and they're they're professionals. Like, this is their job. And what was perfect, I realized, is they tell you in agonizing detail exactly what reasoning they're used to solve those particular puzzles. Right? So we, with their permission, took all of their videos, which represents thousands of hours of very high quality human reasoning, like thought traces, and scraped them and made that available for imitation learning. Right? We did try to do this internally. Turns out that I did a little bit too much of a good job of really creating a very difficult benchmark. Right? So we're still trying to get that stuff working. We'll publish it that if we if we have some success. Yeah. I wanna I wanna really sell the fact that this this reasoning benchmark really is different. Right? Not only do you get something that's super grounded, like, know exactly if it's right or wrong, so you can do RL to your heart's consent, but you can't generalize very easily. Each puzzle is deliberately designed by hand to have a new and unique twist on the rules called a break in that you have to understand. And right now, despite all the progress we've made, the current AI models can't take that leap. They can't find these break ins. Right? They'll fall back to, okay. I'll try no I'll try 5. I'll try 6. I'll try 7. Right? The the reasoning becomes really boring and nothing like what you see in the transcripts that we've we've open sourced from this from this YouTube channel. So I just wanna put the challenge out there. Right? That this this is a a really difficult benchmark, and I think progress on this benchmark will really mean progress in AI generally. Could you reflect a bit? So after watching this cracking the cryptic YouTube channel, how diverse were the patterns? Because Chris was saying to me, oh, you know, these guys, they go on Discord servers. They get these creative crazy ideas. And I'm I'm obsessed. Maybe maybe I'm just being idealistic, but I love this idea of there being a deductive closure of knowledge. Right? That that there's this big tree of of reasoning. And we're all in possession of different parts of the tree to different depths. So the smarter and the more knowledgeable you are, the deeper down the tree you go. But in this idealized form, there is 1 tree. And all knowledge kind of, you know, originates or emanates from these abstract principles. And we could in principle build reasoning engines that could just reason from first principles. And it might be computationally irreducible. So you see, you have to perform all of the steps. And it feels like because we're not in possession of the full tree, what we need to do is kind of fish around. We fish around to find the Lego blocks. Oh, that's a good Lego block. I can apply that to this problem. And maybe that's just what we need to do in AI for the time being is is is we need to just acquire as much of the tree as possible. But could could we just do it all the way down?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence