app / likes
Spark
“Um, the only thing I like out of, out of Codex is the, is like Spark and like yeah.”
Public evidence record
Host · Latent Space
Books, apps, and tools
app / likes
“Um, the only thing I like out of, out of Codex is the, is like Spark and like yeah.”
app / uses
“Yeah. I got, I got the tool. Uh, what, like, I hate, I use Bank of America. I hate bank, I hate the app. Mm-hmm. I hate the web. All banking websites just horrible.”
other / likes
“I like the, the Wiki approach. Uh, my, I’m actually like, uh, you know, obviously I spent some my time at cognition, which, uh, you, you know very well.”
person / uses
“You were very excited because I read Ted Chiang over the holidays and I was very inspired by this short story called Understand, which apparently is, like, pretty old.”
app / uses
“I used to use Overcast. So it would just link to the Overcast page.”
other / uses
“By the way, we use the snip count as a proxy for popularity, right? Because we have download counts, but for example, platforms like Spotify re-host our MP3 file.”
other / recommends
“I think I strongly recommend Jack Bridger's Scaling DevTools, as well as Turner Novak's The Peel.”
other / recommends
“I think I strongly recommend Jack Bridger's Scaling DevTools, as well as Turner Novak's The Peel.”
app / uses
“So that's what basically I use AI News for. I have a lot of prompts and a lot of steps and a lot of criteria and O1 just kind of checks through each kind of systematically.”
tool / built
“It's funny because I built on top of the fork of Bolt.new that already has the multi LLM thing.”
other / uses
“I use XML in other models as well, and it's just a really nice way to make sure that the thing that ends is tied to the thing that starts. That's the only way to do code fences where you're pretty sure example one start, example one end, that is one cohesive unit.”
app / uses
“I used to tell people go to the DevIn demo and look at the four things that they offer and say each of those things is a startup.”
Claim ledger
95 transcript-backed records
01 / evaluation
“I think, uh, I don’t know people say this, but I, I, I don’t think they try it hard.”
02 / evaluation
“Uh, so I think like basically what we’re talking about is the vertical versus horizontal, uh, debate in, in AI startups.”
03 / evaluation
“Minimal in, in a sense of like, the worst you do is you just get hired into one of these labs anyway. So I, I think the, the market for people who just do things and try things and try to execute in like a competent way, even if like it doesn’t work out commercially, even if it just wasn’t that great anyway.”
04 / evaluation
“Uh, one of the reasons I reached out was because you started promoting more sort of internal tooling, uh, primarily Tangle, but also a lot of people have seen and adopted Tobi’s QMD, uh, and obviously, I think, uh, Shopify has always been sort of leading in terms of, uh, engineering.”
05 / evaluation
“Like I think for example, right, like in the audio kind, kind of use cases, the SSMs ef-effectively have unbounded context length because they, they just have to operate on like the most, the sliding window of the most recent stuff.”
06 / evaluation
“Then the other thing that you mentioned, which also raised my eyebrows, was content-based caching, which you mentioned is, is, um, you know, is ve-very much, uh, um, a sort of efficiency measure about, uh, you know, just like recalculation only on, on sort of content addressing Which I think makes sense.”
07 / evaluation
“It’s almost like you’re, it’s like a ratchet. It’s like you’re forcing build time discipline, because if you don’t, it’ll just grow and grow.”
08 / evaluation
“That you guys do. But, I, I could see, I could see that I think the, the human intent is something that people are not even used to because we’re so used to static worlds or, worlds that just don’t react, or, I don’t know.”
09 / evaluation
“Like, I don’t know if, I don’t know if I can say that, but like, you know, um, I think what my point kind of is, is that there’s, like, I look at slopes of the scaling laws and like, this slope is not working, man.”
10 / evaluation
“Yeah I would say, let’s call it a year ago the models weren’t even good enough to do any of this stuff.”
11 / evaluation
“Because Open AI is Codex app doesn’t have a file editor, like it has file viewer, but isn’t a file editor.”
12 / evaluation
“There’s all these like really basic questions that no one stops to answer for people because everyone’s just like too busy launching.”
13 / evaluation
“Yeah. They want to be like we own your logs and give us our, some part of the, [00:17:00] self-healing software that everyone wants.”
14 / evaluation
“I love, I love pushing back. I think that. That is what a lot of technology consultants love to hear this sort of thing, right?”
15 / evaluation
“I, I think you guys have, you know, really made a lot of progress and I think taking a lot of industry leadership for C Bench verified and, and now moving on to C Orange Pro.”
16 / evaluation
“I think motion, you know, I still want to shout out, I think Gemini, still the only native video understanding model that’s out there.”
17 / evaluation
“Totally. And I think that is partially why it made your launch successful because you launch with a sufficiently spanning set of here's examples and then people just copy paste and expand from there.”
18 / evaluation
“I think once you enable two way and once you enable client server to be the same and delegation of work to another MCP server, it's definitely more agentic than not.”
19 / evaluation
“You know, I think you, you maybe have a published some research that says like, actually sometimes to get, to get the model working the right way, you have to do multi-step prompting or jailbreaking to, to, to behave the way that you want.”
20 / evaluation
“Like I, I, to some extent, I think the only reason you and I are talking about it is that they, both of them have reported like ridiculous numbers.”
21 / evaluation
“This is an actual tipping point. And I think I like as people who are like, our function as podcasters and industry analysts is to raise the bar or focus attention on things that you think matter.”
22 / evaluation
“I was gonna say this, so I have a list [00:24:00] of like two years ago we, I wrote the Anatomy of autonomy posts where it was like the, the first, like what's going on in agents and, and and, and, and what is actually making money. Because I think there's a lot of gen I skeptics out there.”
23 / evaluation
“I mean, again, I I, I don't like this label of how Fast Open Source caught up because it's really how Fast Deepsea caught up.”
24 / evaluation
“I would say that, you know, obviously I'm a power user of all these tools. You have done a better job than Descript.”
25 / evaluation
“I mean, on my side, I, I think I watched only like half of the talks. Cause I was running around and I think people saw me like towards the end, I was kind of collapsing.”
26 / evaluation
“I don't know. I think that's, that's me being a slow adopter. No, no. I mean, that's.”
27 / evaluation
“There's a lot of interest. I think Pokemon really is a good agent benchmark, to be honest.”
28 / evaluation
“I think that Swarm has really popularized the handoff technique, which I thought was like, you know, really, really interesting for sort of a multi-agent.”
29 / evaluation
“Yeah, yeah, uh, for sure. And then one thing on the, on like the breadth, you know, I think a lot of the deep research, open deep research implementations have this sort of hyper parameter about, you know, how deep they're searching and how wide they're searching.”
30 / evaluation
“Where he basically observed that the browser is turning the operating system into a poorly debugged set of device drivers, because most of the apps are moved from the OS to the browser.”
31 / evaluation
“Just because like you're one of the, you know, best performing, I think, LLM tool companies that have started up in the last couple of years.”
32 / evaluation
“I would say I was very keen on, I think even at the end of last year, people were already saying it was one of the most exciting agents that was coming out of Google.”
33 / evaluation
“I'm calling this out because everyone is like, oh my God, it takes hours for, it does hours of work autonomously for me.”
34 / evaluation
“That was super counterintuitive for us. So actually, the first time I realized that, what you're saying is when I was talking to Jason Calacanis and he was like, do you actually just make the answer in 10 seconds and just make me wait for the balance?”
35 / evaluation
“Yep. Because I think at that point, like, users will just drop off. Nope. But what's been surprising is, like, that's not the case at all.”
36 / evaluation
“Drafting anything like I want to draft like copy for my conference that I'm running, like I'll put it there first and then I like, it'll just have the canvas up and I'll just say what I don't like about it and it changes.”
37 / evaluation
“I think because it's maybe sold as like sort of writing help when really like it's kind of, it's the scratch pad.”
38 / evaluation
“I think the naming as well matters. It seemed like a branch off of the main, main tree of development.”
39 / evaluation
“Small and yeah. Because it used to be latent diffusion models and then they trained it up.”
40 / evaluation
“As someone who also does charts, XAI is continually snubbed because they don't work well with the benchmarking people.”
41 / evaluation
“I think the quote that I highlighted in AI News was that it is the best, like Blackwell is the best selling series.”
42 / evaluation
“I think the Researchers that I talked with at NeurIPS were kind of positive on this because basically you need private test [00:14:00] sets to prevent contamination.”
43 / evaluation
“This one was more like the money one which you know it's funny because I think developers are like quite uninterested in money.”
44 / evaluation
“Noam and basically everyone on the Strawberry team was very insistent that what they did for reinforcement learning, chain of thought, cannot be replicated by a whole bunch of open source model calls. Do you think that that is wrong?”
45 / evaluation
“And then, do you see any more opportunity on the... You know, I think you made a big splash with 1,000 tokens per second.”
46 / evaluation
“Cosign was doing well on SweetBench, but they didn't want to leak those results. So that's why you don't see O1 preview on SweetBench, because they don't submit their reasoning choices.”
47 / evaluation
“Basically it's just like the meta version of whatever Hugging Face offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is. But to me, it's not clear that just because Meta open sources Lama, that the rest of LamaStack will be adopted.”
48 / evaluation
“The LoRa thing is interesting because I think you also, the reason people add additional costs to it, it's not because they feel like charging people.”
49 / evaluation
“Because if you go from one building to two buildings, congrats, you're now remote from the other building.”
50 / evaluation
“The framing can be different if you were, so I think tinkerers has this connotation of not serious or like small.”
51 / evaluation
“I think that's good. I'll maybe carve out that I think the UK has done really well.”
52 / evaluation
“This is something I think about for AI engineering as well, which is the big labs want you to hand over everything in the prompts, and only code of English, and then the smaller brains, the GPU pours, always want to write more code to make things more deterministic and reliable and controllable.”
53 / evaluation
“that are starting to provide integrations as a service, right? I used to work in an integrations company.”
54 / evaluation
“There's all these other companies that are like, we will do the integrations for you.”
55 / evaluation
“It's funny because in some ways, the model labs are competing for you, right? You don't have to do any effort.”
56 / evaluation
“I think my fundamental philosophical doubt is, does the router model have to be at least as smart as the smartest model?”
57 / evaluation
“The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.”
58 / evaluation
“Top P top cake. It doesn't matter. I'll just put it in the docs and you figure it out.”
59 / evaluation
“I will say I've been dealing with EDB a little bit from my conference, and they've been extremely responsive and it's been nice to see, because I never get to see this out of government, nice to see that as someone that wants to bring a foreign business into Singapore, they're kind of rolling on the welcome mat.”
60 / evaluation
“I'm pretty science-based, like, you know, but probably the most like spiritual woo-woo thing about me is I don't think that would lead to consciousness or AGI just because like there's something in- there's a soul, you know?”
61 / evaluation
“I think OpenAI was just like, look, this is a thing now. We have to fix this. These students just rushed it.”
62 / evaluation
“You're one of like, you're maybe the first PhD thesis defense I've ever watched in like this AI world, because most people just publish single papers, but every paper of yours is a banger.”
63 / evaluation
“I will mostly agree and I'll slightly disagree in terms of this, which is like, whether designing for humans also overlaps with designing for AI. So Malte Ubo, who's the CTO of Vercel, who is creating basically JavaScript's competitor to LangChain, they're observing that basically, like if the API is easy to understand for humans, it's actually much easier to understand for LLMs, for example, because they're not overloaded functions.”
64 / evaluation
“Like, I actually have most of my problems with AI news when the model thinks it knows more than it knows because it combines knowledge with intelligence.”
65 / evaluation
“I actually might tweak my approach based on that, because I was trying to give bad examples of do not do this, and it still does it, and maybe that doesn't work.”
66 / evaluation
“Endorse all that. And I think getting things into structured output and doing those scoring is a very core part of AI engineering that we don't talk about enough.”
67 / evaluation
“You've done very well. And I think you've honestly done the community a service by reading all these papers so that we don't have to, because the joke is often that, you know, what is one prompt is like then inflated into like a 10 page PDF that's posted on archive.”
68 / evaluation
“I don't know what to make about it because I don't think it's adopted seriously by the large labs.”
69 / evaluation
“Whenever image generation is concerned, obviously because of the Gemini issue, it is very tricky for large companies to release that.”
70 / evaluation
“These are the same things to the model. That is a huge, huge win for interpretability, because up to now, we were only doing interpretability on toy models, like a few million parameters, a model of Go or chess or whatever.”
71 / evaluation
“We love just Meta's commitment to open source and, you know, you do what you need to do to make it work for your organization.”
72 / evaluation
“The other thing is RLHF came from the alignment community. And I think there's a lot of conception that maybe it's due to safety concerns, but I feel like it's really over the past two, three years expanded to just this produces a better model period, even if you don't really are not that concerned about existential risk.”
73 / evaluation
“We're basically saying that, I think that one of the lessons from AlphaGo is that people thought that human interest in Go would be diminished because computers are better than humans.”
74 / evaluation
“I don't know how to describe, you've done so much work in a very short amount of time at Meta, but you were most notably leading Llama 2 and now today we're also coordinating on the release of Llama 3.”
75 / evaluation
“Because I didn't know that, I don't know how much to believe, you know, like there's a lot of these kinds of papers where it makes a lot of noise, but it doesn't actually pan out.”
76 / evaluation
“I think, yeah, very underrated, very underrated, this sort of PhD with industry expertise, because you're also publishing papers the whole time.”
77 / evaluation
“Well, I think your progress into NLP was like really strong, because like the first thing you worked on at Meta was Bloom.”
78 / evaluation
“I think obviously MMLU Pro is the top one, just because that's the top number that a lot of people report.”
79 / evaluation
“Yeah, so I really like this concept of evaluation. So actually, yeah, I think there's typically what I always say is like sort of 25 is random chance, 50 is average human, 75 is expert human, 90 is you're cheating.”
80 / evaluation
“Obviously, I think you're like our second or third person from Hugging Face on the podcast and it's like the definitional sort of open AI company, maybe the real open AI.”
81 / evaluation
“Like, just because you re parameterize some benchmarks in evals and, you know, make it linear, doesn't mean emergence is completely gone.”
82 / evaluation
“Like, or So it's a very interesting observation where like, most efficiency work is just busy work, or like, it's work at a small scale that doesn't, that just ignores the fact that like, this thing doesn't scale, because you haven't scaled it.”
83 / evaluation
“I think a lot of people are exploring that and I think every now and then people get a bout of knowledge graph religion and then it kind of doesn't work out.”
84 / evaluation
“One thing I'll, one thing I'll mention quickly is that a lot of the stuff that you mentioned is typically not part of the normal interview loop. It's actually really hard to interview for because this is the stuff that you polish out in, as you go into production, the coding interviews are typically about the happy path.”
85 / evaluation
“I have some appreciation, I think when you had me on your podcast, I was still working at Temporal and that was like a nice Framework, if you live within Temporal's boundaries, you can pretend that all those faults don't exist, and you can, you can code in a sort of very fault tolerant way.”
86 / evaluation
“If some AI engineer would not know, I don't know what, , I don't know where we would stoop to, to call something required knowledge, , or you're not part of the cool kids club.”
87 / evaluation
“The ML first mindset, I think, is something that I struggle with as well, because the errors, when they do happen, are bad.”
88 / evaluation
“I, I do often say that I think AI engineering is about 90 percent software engineering with like the, the 10 percent of like really strong really differentiated AI engineering.”
89 / evaluation
“One thing I see in the recent papers that have been coming out is this sort of concept of multi-stage training data. And if you're doing full fine tuning, maybe the move or the answer is don't train 500 billion tokens on just code, because then yeah, it's going to massively overfit to just code.”
90 / evaluation
“Yeah, I think, you know, the one thing that makes this sort of generative AI era very different from the sort of data science-y type era is that it is very non-deterministic and it's hard to control.”
91 / evaluation
“It's just retrieval. And here it's like, The home of generative AI, this, whatever hyperstition is in my mind, like this is actually pushing the edge of what generative and creativity in AI means.”
92 / evaluation
“When you say you need a certain set of tools for people to sort of invent things from first principles Devin is the agent that I think has been able to utilize its tools very effectively.”
93 / evaluation
“I think one of the agents loopholes or one of the things that is a real barrier for agents is LLMs really like to get stuck into a lane.”
94 / evaluation
“Because I just didn't find them reliable because they just hallucinated their own uncertainty.”
95 / evaluation
“Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.”