High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

6 Jun 2026 The Cognitive Revolution AI in the AM — Week 1 Highlights (June 2026)

“They can't solve most of those, whereas you get those real human scientists and you get most of them. So the summary of that is it's really easy to be excited when they work well, but you got to pay attention to all the really simple ways that they break before you get too excited, I think.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
6 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…It surprises me and doesn't surprise me all at the same time. Some days I wake up and I feel like I'm living in the future, and other days I wake up and I feel like I'm, living in this sort of strange reality with all these agents that can't do the things that I want them to do. I'll give an example that's really grounded in AI and scientific discovery. So we have this project Code Scientist, which looks very similar to a lot of the projects that sort of you pull up on Twitter every day. where people say, I made this AI agent and which is like a thin wrapper on some open AI or cloud model or whatnot. And it generates code automatically and it generates ideas automatically and runs in a loop and away they go, writes papers. And so we gave it 50 research ideas and let it chug away for a couple of days. And after a few days it came back and it said, I've discovered 19 new things, and we were very excited, right? Wow, 19 new things, we live in the future, life is great, all that jazz. And so I wrote papers on those 19 new things, and we gave those 19 papers to three colleagues at AI2 who hadn't seen it before and said, tell me if this is a real discovery, you know, look through these papers, here we go. And they went through them, and I think it was 70 or 80% of the papers. I said, yeah, it's probably at least incrementally novel and minimally scientifically sound and whatnot. And then we convinced somebody, me, to go through and spend days and days and days looking at the thousands upon thousands of lines of code that these models were generating to support their discoveries. And it went down to like 30% of the discoveries were probably real. And the things that you see are absolutely all over the place. One fun example is the AI came up with some fancy idea for making a new neural network architecture with some fancy new kind of attention. I don't know. It wrote hundreds of lines of Python code with all this neural network code that I have absolutely no idea what it was doing or can't understand any of it. And so I'm going through and I'm like, how on earth am I going to review this? is in my domain area. And then I get to like the end of a couple 100 lines of code. And there's just this comment that says, comment, insert rest of neural network code here. And then it picked a random number and returned a random number from that function. And so this model, this paper, this entire paper was analyzing the values of a random number generator. And that, you know, isn't shown to the, you know, nobody knows that if you're reading the paper. But the science itself, it's hard to evaluate it. It's hard to be sound. And so a lot of this, when you see it do something amazing, it's easy to be very impressed. But then when you use a standard benchmark, like we have Science World and Discovery World, these sort of virtual environment benchmarks, Science World does 4th grade science, Discovery World does sort of like a master's or PhD level science. , like we have Science World and Discovery World, these sort of virtual environment benchmarks, Science World does 4th grade science, Discovery World does sort of like a master's or PhD level science. The best models right now are getting something like 80% on the 4th grade science. So you ask them to go in this environment and boil water and they can't do it 20% of the time, right? That's wild. Or you ask them to go in and give them a toy task, you know, the colonists on Planet X are getting sick. figure out why and solve it. They're really terrible at that, right? They can't solve most of those, whereas you get those real human scientists and you get most of them. So the summary of that is it's really easy to be excited when they work well, but you got to pay attention to all the really simple ways that they break before you get too excited, I think. That's not to say they don't have utility and there's lots of places that they do have Lots of very near-term utility, but I think my job is sick for a little. Now, limits like that do cap how far you can lean on these systems today. But my basic read is pretty simple. They can now do a great many things, super reliably, that they used to be terrible at. And even a 30% real discovery rate is hard to understand as anything but the beginning of a science fiction future. The clearest stakes this week were in security. The best guest we had breaks into companies for a living. And his read on where AI actually bites surprised me.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence