High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Thomas Ahle: evaluation

28 Jun 2026 Machine Learning Street Talk The Thermodynamic AI Computing Chip - Thomas Ahle

“Like are you when you do reinforcement learning, I mean the chain of thought before the reinforcement learning and after reinforcement learning I think is very different because before it was kind of you try and prompt it and some tricks they kind of work and kind of doesn't work.”

— Thomas Ahle

Source trail

Everything needed to verify it.

Speaker
Thomas Ahle
Attribution
Verified speaker
Claim type
evaluation
Recorded
28 Jun 2026
Publisher
Machine Learning Street Talk

Transcript context

…How much can we read into chain of thought? Right? Because you because, you know, like, some people just call it chain of thoughtlessness, like, you know, and, but actually, it is probably the the modus operandi now for doing interpretability. Right? For actually understanding what they're what they're thinking. Yeah. But, you know, you you could argue that chain of thought is like the press secretary, not the orchestrator. So it's it's almost a post hoc confabulation. And but that does that's not quite right, is it? Because that sounds a little bit like there's no cause or link between the chain of thought and what the language model outputs. That's not true. So how much can we read into it? Yeah. It's like you want to give the model somewhere to think, right? Like are you when you do reinforcement learning, I mean the chain of thought before the reinforcement learning and after reinforcement learning I think is very different because before it was kind of you try and prompt it and some tricks they kind of work and kind of doesn't work. But once you do the reinforcement learning you need like, it's like a Turing machine, right? Like it needs to have infinite memory and being able to have something to operate on. Whereas when you just had the transformer and just like single shot output, it's like, I mean probably in the Chomsky hierarchy it would be a completely different type of system like, I don't know, an atomic or something, right? There's like only finite amount of computation it can do, but now it can do as much computation as it wants. And it just has to learn to figure out how to do it. Like you could imagine building a really simple LLM type system and with access to chain of thought it would be a universal Turing machine. So now at that point it's just now because it can have this tape and it can keep reading and putting. I think I think that would be pretty easy to make a structure like that. But then the question is whether it can learn it. But at least now it has the like representation capacity so it can do it and clearly something is improving and working. This is actually the bigger problem, like, with the ecosystem, which is that software was always designed to be a thing that that reduces complexity and introduces canalization. So spreadsheets, they're they're a great example of this. Now, you know, accountants use spreadsheets and financial and, you know, like, everyone's using spreadsheets, and it creates an interface that everyone uses, and it reduces complexity in the system. This this agentic AI, it just creates spaghetti everywhere. And this is part of the reason why people aren't shipping because, you know, it creates ephemeral, you know, complex it it's a little bit like bash scripting on steroids. So I've now created this web that only I understand and it's becoming more and more specialized over time, I can't share it with other people. And and the entire ecosystem is becoming very messy. Yeah. No. I I definitely think…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence