person / recommends
Kevin Roose
“On The New York Times, I, I share your disappointment. I, I recommend Kevin Roose and, and Hard Fork out of The New York Times family of-”
Public evidence record
Host · The Cognitive Revolution
Books, apps, and tools
person / recommends
“On The New York Times, I, I share your disappointment. I, I recommend Kevin Roose and, and Hard Fork out of The New York Times family of-”
tool / uses
“And certainly if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount. What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
tool / uses
“And certainly if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount. What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
app / uses
“What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
app / uses
“What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
tool / uses
“I think I used Gemini 3 flash at the time.”
app / uses
“That's probably the biggest reason that I use Claude is that I perceive it to be most robust to that kind of stuff.”
app / uses
“I use Beeper Desktop to try to like aggregate a half dozen or so of them. That's also kind of painful. I feel like Beeper Desktop is, Good idea crashes a lot for me.”
other / likes
“Is there stuff that we can do or is it, you know, is there is it a different organization's job to figure out how to fill that gap. Because I do feel like I want some more, and I love some of the AE Studio stuff, including self-other overlap.”
Claim ledger
44 transcript-backed records
01 / evaluation
“am a little more sympathetic, I think, off the top to the argument that there was some massive reappropriation of human knowledge that is upstream of all AI.”
02 / evaluation
“Definitely getting bogged down sometimes in listening to all these Suno song generations and trying to iterate to find something that I actually feel like I really like.”
03 / evaluation
“In today's world, one of the reasons I can't send my AIS out to do all my stuff for me is that humans are pretty clever about tricking and ripping off the AIS, so I'm not sure how we avoid a situation.”
04 / evaluation
“I think if you survey most people that use AIA lot, they would say, Oh yeah, Claude's the most aligned right.”
05 / evaluation
“Forecasting gives us an opportunity to do some world modelling. So Future Search talked about this a little bit at the Manifest conference a couple weeks ago and the feature in the product is rolling out I think literally today.”
06 / evaluation
“The iteration time from model to model is now potentially shorter than the time horizon that it would take a model to top out in terms of the absolute best performance on a super hard, ambitious, you know, long running task. So I, I had even heard him kind of propose something along the lines of like a claw back or sort of a recall program almost where, and obviously this doesn't work in open source, but it can work in AAPI paradigm where a model might get released, you know, day N after it's been deemed to be ready.”
07 / evaluation
“So there's like a really lot of unpack there here. So just like a little bit of the background, the reason that we started to create our own foundation on models like this realisation that what closed model providers are offering does not make sense for us economically.”
08 / evaluation
“They basically think that what they're doing is somewhat near optimal and any sort of accuracy improvements you're going to get over them is going to be tiny and like hard to understand. And I think that's just because we only really understand human intelligence.”
09 / evaluation
“Because the ablation of the J space just leaves causes such a performance degradation on these like hard multi step type of tasks that if you don't see concepts in the J space, you can be, they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etcetera, etcetera.”
10 / evaluation
“Because like you, the sensor dynamic range is limited and then you're losing either some details in highlights or in shadows or for example, let's say you're taking a stream from a camera and want to simulate how it looks like with a different focal length.”
11 / evaluation
“And I tried my best over basically like, you know, 12 to 16 hours of the Fable situation. I think I made a pretty good model.”
12 / evaluation
“Overall, Pangram is quite accurate and yet we have at least one example out of 400 or so essays where I think the zero score I would confidently assert is wrong and unfair and should not be the basis for like a pylon.”
13 / evaluation
“This is very useful for us because we can evaluate things immediately. So when Fable came out, the first time the clawed Fable came out, we were able to evaluate it within 24 hours and it was the best single agent forecaster on our leaderboard.”
14 / evaluation
“Now, excuse me, what I What sort of jumped out at me in terms of your approach is that you've developed a architecture search process where the promise to customers is not that, hey, we developed this one paradigm and the old calculus teacher used to say, when all you have is a hammer, everything looks like a nail.”
15 / evaluation
“Back in 2023, he called generative AI an existential threat, put his entire org on AI one day a week, and when most of his people pushed back, he replaced them, rebuilding around what he calls AI DNA. We started with one of his recent acquisitions, a company called Chorus.”
16 / evaluation
“I think that is really underappreciated by the public at large. It's more appreciated by the people developing the AIs because they've at least had an experience that was formative for me when I was doing the GPT-4 Red Team close to four years ago now.”
17 / evaluation
“If I try to channel Balaji for a second, which I wouldn't pretend to be able to do it an A+ job of, I think he would say something like, "We all have way too much faith in the U- US government.”
18 / evaluation
“You know, the responses have been, "Well, you know, the ultrasound doesn't see this that well, doesn't see that well," or, you know, "We've, we don't actually recommend whole body scans because, you know, there's a lot of false positives," and all this kind of stuff.”
19 / evaluation
“The other thing that's kind of related to this that jumped out at me is a sort of escalation, I guess, of both the difficulty of monitoring and some recent advances in monitoring techniques that I'm not sure exactly where they leave us on net. But we both see in the system card examples of extremely illegible chain of thought, which, you know, is just like this wall of emojis and sort of, you know, non-human language symbols strung together that I think is pretty spooky and, like, definitely, um, you know, don't like to see that, to put it simply and mildly.”
20 / evaluation
“I think people like me sort of run a bit of a risk of getting detached from, especially because I work by myself largely these days, kind of run a risk of getting detached from what's going on in the real world at real companies that are actually driving most of the economy and where not everybody has the luxury or the inclination to be a bleeding edge early adopter with all of the, I'd say more ups than downs, certainly, but certainly a mix of ups and downs that come with that.”
21 / evaluation
“A company running thin margins on top of Opus is going to struggle to say, no, don't use that.”
22 / evaluation
“If you believe that models have their own deep-seated goals and that those goals might diverge from ours, then this could be very bad, right? It could be like, it could be extremely bad because they would be using this reasoning to figure out how to please us while like still having their own goals.”
23 / evaluation
“Do they handle the sort of confirmation step well? Because I one thing that was flagged for me as I was talking to AI, of course, about how to do this is that the, those sort of VoIP numbers sometimes don't work for like, you know, you sign up for a new account, then you get the, the code or whatever.”
24 / evaluation
“I want to get into that and get your take on how people who maybe don't work directly in the space yet or who do what kind of work but are not sure if they're making the biggest impact that they can, how they might think about pivoting their careers to try to have the most positive, try to make the most positive contribution that they can.”
25 / evaluation
“AI is one of the premises of this show and one of the reasons I enjoy making it so much is that it's obviously a general purpose technology, a horizontal layer, something that kind of intersects with everything.”
26 / evaluation
“I mean, you can maybe interpret inoculation prompting differently than I will, but my general description of inoculation prompting is there's a generalization, a very problematic generalization that happens if you reward the model during reinforcement learning for something you didn't quite intend, especially if it's like a flagrant hack, then the model can sort of start to generalize to I'm the kind of thing that loves to reward hack and I get rewarded for that.”
27 / evaluation
“I think a general sketch would be like 03 might be the most misaligned model that was ever released to the public. It seemed like it was right in that tween zone where RL had really scaled up and some of these problems were starting to show up, and since then there's been a bunch of work to try to reduce them.”
28 / evaluation
“I do know that they have to be a lot faster because the ad's gotta show up really quickly on the page. And then I know also that there's a pretty challenging matching problem in there somewhere because I've got millions of, you've got, we've got, society collectively has got millions of these profiles of individuals.”
29 / evaluation
“one thing I will say about AI is it is allowing me to create stuff that I don't think is terrible, at least, and that I enjoy the process of creating in ways that I just never would have had any opportunity to do before.”
30 / evaluation
“When you describe, you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be to have any sort of advantage.”
31 / evaluation
“I have AI, have my own test. And it's been, you know, every time we try it, it fails and I'm like, OK, another, another one doesn't work right.”
32 / evaluation
“Try doing bunch of these things that like you wouldn't want someone participating in like the water economy to do because and I think quite a lot of these things it's like illegal, like price collusion and stuff like this.”
33 / evaluation
“Because I, I often feel like you almost, you're almost kind of trying to prompt inject the LLM which is running the search and you're trying to get in there and hack it so that your, you know, your page goes up.”
34 / evaluation
“It, it strikes me that we haven't really seen the true unleashing of the Internet's adversarial potential. And so, you know, that's one thing that they, I, I would say one of their biggest weaknesses, even Frontier models biggest weaknesses these days is how gullible they remain.”
35 / evaluation
“Now, the problem is today we don't know how to build abstractions in a robust and scalable way that, you know, sort of represent the noise of that underlying substrate.”
36 / evaluation
“It turns out for the kinds of capacitors we use, you see variations that are on the order of, you know, sort of 10 parts per million, right. So giving you levels of precision that are in the neighborhood of 20 bits of precision, which it turns out is well beyond what we need for the quantization kinds of, you know, levels that we care about, which are typically at the level of eight bits and you know, higher than that in some cases.”
37 / evaluation
“Turns out we don't need anywhere near that precision for the capacitors that we use. But, but it's really because of this alignment with this geometric control that this particular approach has that allows it to be brutally accurate in the ways that you need it to be through all of these layers of abstraction to be able to scale up.”
38 / evaluation
“I don't know, maybe simplifying oversimplifying this a bit, but interventions of that sort seem maybe not any or all, but like in general seem to promote affirmative responses from models such that maybe you could say, you could, you know, once you make these kind of interventions, they'll say yes to anything.”
39 / evaluation
“I, you know, I think we talked with this more last time than this time, but this notion of mutualism as a positive vision for the future, I think is another major strength of just everything that you bring to the table.”
40 / evaluation
“It's got like my Claude MD and it's got access to like my, you know, sort of who Nathan is and all the, you know, I'm building up a lot of context that it has consistent access to every time. So I think in that sense, like I sort of see this like whole model versus, you know, single thread thing as kind of being blurred anyway, because I've got the same like rather large prompt that I'm using every time.”
41 / evaluation
“Obviously, people have radically different understandings of what's coming, everything from still outright denialism, which I think is increasingly discredited and can be ignored, but there's still this sort of more credible version of AI as normal technology.”
42 / evaluation
“If I think, though, even just about my own ability to search through my own stuff, my own Gmail, my own Google Docs, One of the intuitions I have pretty strongly is if I were to give you full access to my Gmail and give you full access to my Google Docs, you couldn't search through it nearly as well as I can.”
43 / evaluation
“Yeah, so that brings up another, I think, huge question for AI safety research in general, and probably the strongest, maybe not in, I don't know if you would say strongest in the sense of being most compelling to you, but certainly the most hawkish or fiercest criticism that AI safety research gets is that it always ends up being dual use and that it always ends up somehow accelerating the core capabilities track.”
44 / evaluation
“Then now obviously we've got pretty amazing language models, I would imagine that like the best language models are maybe an overkill for some of the use cases, if only because of cost and latency.”