High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Shawn Wang

Host · Latent Space

Claims
500
Episodes
64
Shows
1
Named items
20

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

app / likes

Spark

“Um, the only thing I like out of, out of Codex is the, is like Spark and like yeah.”

Latent Space · 23 Apr 2026

Evidence receipt · Source ↗

app / uses

Bank of America

“Yeah. I got, I got the tool. Uh, what, like, I hate, I use Bank of America. I hate bank, I hate the app. Mm-hmm. I hate the web. All banking websites just horrible.”

Latent Space · 20 Mar 2026

Evidence receipt · Source ↗

other / likes

Wiki approach

“I like the, the Wiki approach. Uh, my, I’m actually like, uh, you know, obviously I spent some my time at cognition, which, uh, you, you know very well.”

Latent Space · 5 Mar 2026

Evidence receipt · Source ↗

person / uses

Ted Chiang

“You were very excited because I read Ted Chiang over the holidays and I was very inspired by this short story called Understand, which apparently is, like, pretty old.”

Latent Space · 6 Feb 2026

Evidence receipt · Source ↗

app / uses

Overcast

“I used to use Overcast. So it would just link to the Overcast page.”

Latent Space · 14 Mar 2025

Evidence receipt · Source ↗

other / uses

Snip

“By the way, we use the snip count as a proxy for popularity, right? Because we have download counts, but for example, platforms like Spotify re-host our MP3 file.”

Latent Space · 14 Mar 2025

Evidence receipt · Source ↗

other / recommends

The Peel

“I think I strongly recommend Jack Bridger's Scaling DevTools, as well as Turner Novak's The Peel.”

Latent Space · 28 Feb 2025

Evidence receipt · Source ↗

other / recommends

Scaling DevTools

“I think I strongly recommend Jack Bridger's Scaling DevTools, as well as Turner Novak's The Peel.”

Latent Space · 28 Feb 2025

Evidence receipt · Source ↗

app / uses

AI News

“So that's what basically I use AI News for. I have a lot of prompts and a lot of steps and a lot of criteria and O1 just kind of checks through each kind of systematically.”

Latent Space · 1 Feb 2025

Evidence receipt · Source ↗

tool / built

Bolt.new

“It's funny because I built on top of the fork of Bolt.new that already has the multi LLM thing.”

Latent Space · 2 Dec 2024

Evidence receipt · Source ↗

other / uses

XML

“I use XML in other models as well, and it's just a really nice way to make sure that the thing that ends is tied to the thing that starts. That's the only way to do code fences where you're pretty sure example one start, example one end, that is one cohesive unit.”

Latent Space · 28 Nov 2024

Evidence receipt · Source ↗

app / uses

DevIn

“I used to tell people go to the DevIn demo and look at the four things that they offer and say each of those things is a startup.”

Latent Space · 2 Aug 2024

Evidence receipt · Source ↗

Claim ledger

What Shawn said.

95 transcript-backed records

03 / evaluation

Minimal in, in a sense of like, the worst you do is you just get hired into one of these labs anyway. So I, I think the, the market for people who just do things and try things and try to execute in like a competent way, even if like it doesn’t work out commercially, even if it just wasn’t that great anyway.

“Minimal in, in a sense of like, the worst you do is you just get hired into one of these labs anyway. So I, I think the, the market for people who just do things and try things and try to execute in like a competent way, even if like it doesn’t work out commercially, even if it just wasn’t that great anyway.”
Speaker
Shawn Wang
Publisher
Latent Space

04 / evaluation

Uh, one of the reasons I reached out was because you started promoting more sort of internal tooling, uh, primarily Tangle, but also a lot of people have seen and adopted Tobi’s QMD, uh, and obviously, I think, uh, Shopify has always been sort of leading in terms of, uh, engineering.

“Uh, one of the reasons I reached out was because you started promoting more sort of internal tooling, uh, primarily Tangle, but also a lot of people have seen and adopted Tobi’s QMD, uh, and obviously, I think, uh, Shopify has always been sort of leading in terms of, uh, engineering.”
Speaker
Shawn Wang
Publisher
Latent Space

05 / evaluation

Like I think for example, right, like in the audio kind, kind of use cases, the SSMs ef-effectively have unbounded context length because they, they just have to operate on like the most, the sliding window of the most recent stuff.

“Like I think for example, right, like in the audio kind, kind of use cases, the SSMs ef-effectively have unbounded context length because they, they just have to operate on like the most, the sliding window of the most recent stuff.”
Speaker
Shawn Wang
Publisher
Latent Space

06 / evaluation

Then the other thing that you mentioned, which also raised my eyebrows, was content-based caching, which you mentioned is, is, um, you know, is ve-very much, uh, um, a sort of efficiency measure about, uh, you know, just like recalculation only on, on sort of content addressing Which I think makes sense.

“Then the other thing that you mentioned, which also raised my eyebrows, was content-based caching, which you mentioned is, is, um, you know, is ve-very much, uh, um, a sort of efficiency measure about, uh, you know, just like recalculation only on, on sort of content addressing Which I think makes sense.”
Speaker
Shawn Wang
Publisher
Latent Space

07 / evaluation

It’s almost like you’re, it’s like a ratchet. It’s like you’re forcing build time discipline, because if you don’t, it’ll just grow and grow.

“It’s almost like you’re, it’s like a ratchet. It’s like you’re forcing build time discipline, because if you don’t, it’ll just grow and grow.”
Speaker
Shawn Wang
Publisher
Latent Space

08 / evaluation

That you guys do. But, I, I could see, I could see that I think the, the human intent is something that people are not even used to because we’re so used to static worlds or, worlds that just don’t react, or, I don’t know.

“That you guys do. But, I, I could see, I could see that I think the, the human intent is something that people are not even used to because we’re so used to static worlds or, worlds that just don’t react, or, I don’t know.”
Speaker
Shawn Wang
Publisher
Latent Space

09 / evaluation

Like, I don’t know if, I don’t know if I can say that, but like, you know, um, I think what my point kind of is, is that there’s, like, I look at slopes of the scaling laws and like, this slope is not working, man.

“Like, I don’t know if, I don’t know if I can say that, but like, you know, um, I think what my point kind of is, is that there’s, like, I look at slopes of the scaling laws and like, this slope is not working, man.”
Speaker
Shawn Wang
Publisher
Latent Space

15 / evaluation

I, I think you guys have, you know, really made a lot of progress and I think taking a lot of industry leadership for C Bench verified and, and now moving on to C Orange Pro.

“I, I think you guys have, you know, really made a lot of progress and I think taking a lot of industry leadership for C Bench verified and, and now moving on to C Orange Pro.”
Speaker
Shawn Wang
Publisher
Latent Space

17 / evaluation

Totally. And I think that is partially why it made your launch successful because you launch with a sufficiently spanning set of here's examples and then people just copy paste and expand from there.

“Totally. And I think that is partially why it made your launch successful because you launch with a sufficiently spanning set of here's examples and then people just copy paste and expand from there.”
Speaker
Shawn Wang
Publisher
Latent Space

19 / evaluation

You know, I think you, you maybe have a published some research that says like, actually sometimes to get, to get the model working the right way, you have to do multi-step prompting or jailbreaking to, to, to behave the way that you want.

“You know, I think you, you maybe have a published some research that says like, actually sometimes to get, to get the model working the right way, you have to do multi-step prompting or jailbreaking to, to, to behave the way that you want.”
Speaker
Shawn Wang
Publisher
Latent Space

21 / evaluation

This is an actual tipping point. And I think I like as people who are like, our function as podcasters and industry analysts is to raise the bar or focus attention on things that you think matter.

“This is an actual tipping point. And I think I like as people who are like, our function as podcasters and industry analysts is to raise the bar or focus attention on things that you think matter.”
Speaker
Shawn Wang
Publisher
Latent Space

22 / evaluation

I was gonna say this, so I have a list [00:24:00] of like two years ago we, I wrote the Anatomy of autonomy posts where it was like the, the first, like what's going on in agents and, and and, and, and what is actually making money. Because I think there's a lot of gen I skeptics out there.

“I was gonna say this, so I have a list [00:24:00] of like two years ago we, I wrote the Anatomy of autonomy posts where it was like the, the first, like what's going on in agents and, and and, and, and what is actually making money. Because I think there's a lot of gen I skeptics out there.”
Speaker
Shawn Wang
Publisher
Latent Space

25 / evaluation

I mean, on my side, I, I think I watched only like half of the talks. Cause I was running around and I think people saw me like towards the end, I was kind of collapsing.

“I mean, on my side, I, I think I watched only like half of the talks. Cause I was running around and I think people saw me like towards the end, I was kind of collapsing.”
Speaker
Shawn Wang
Publisher
Latent Space

29 / evaluation

Yeah, yeah, uh, for sure. And then one thing on the, on like the breadth, you know, I think a lot of the deep research, open deep research implementations have this sort of hyper parameter about, you know, how deep they're searching and how wide they're searching.

“Yeah, yeah, uh, for sure. And then one thing on the, on like the breadth, you know, I think a lot of the deep research, open deep research implementations have this sort of hyper parameter about, you know, how deep they're searching and how wide they're searching.”
Speaker
Shawn Wang
Publisher
Latent Space

30 / evaluation

Where he basically observed that the browser is turning the operating system into a poorly debugged set of device drivers, because most of the apps are moved from the OS to the browser.

“Where he basically observed that the browser is turning the operating system into a poorly debugged set of device drivers, because most of the apps are moved from the OS to the browser.”
Speaker
Shawn Wang
Publisher
Latent Space

34 / evaluation

That was super counterintuitive for us. So actually, the first time I realized that, what you're saying is when I was talking to Jason Calacanis and he was like, do you actually just make the answer in 10 seconds and just make me wait for the balance?

“That was super counterintuitive for us. So actually, the first time I realized that, what you're saying is when I was talking to Jason Calacanis and he was like, do you actually just make the answer in 10 seconds and just make me wait for the balance?”
Speaker
Shawn Wang
Publisher
Latent Space

36 / evaluation

Drafting anything like I want to draft like copy for my conference that I'm running, like I'll put it there first and then I like, it'll just have the canvas up and I'll just say what I don't like about it and it changes.

“Drafting anything like I want to draft like copy for my conference that I'm running, like I'll put it there first and then I like, it'll just have the canvas up and I'll just say what I don't like about it and it changes.”
Speaker
Shawn Wang
Publisher
Latent Space

44 / evaluation

Noam and basically everyone on the Strawberry team was very insistent that what they did for reinforcement learning, chain of thought, cannot be replicated by a whole bunch of open source model calls. Do you think that that is wrong?

“Noam and basically everyone on the Strawberry team was very insistent that what they did for reinforcement learning, chain of thought, cannot be replicated by a whole bunch of open source model calls. Do you think that that is wrong?”
Speaker
Shawn Wang
Publisher
Latent Space

46 / evaluation

Cosign was doing well on SweetBench, but they didn't want to leak those results. So that's why you don't see O1 preview on SweetBench, because they don't submit their reasoning choices.

“Cosign was doing well on SweetBench, but they didn't want to leak those results. So that's why you don't see O1 preview on SweetBench, because they don't submit their reasoning choices.”
Speaker
Shawn Wang
Publisher
Latent Space

47 / evaluation

Basically it's just like the meta version of whatever Hugging Face offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is. But to me, it's not clear that just because Meta open sources Lama, that the rest of LamaStack will be adopted.

“Basically it's just like the meta version of whatever Hugging Face offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is. But to me, it's not clear that just because Meta open sources Lama, that the rest of LamaStack will be adopted.”
Speaker
Shawn Wang
Publisher
Latent Space

52 / evaluation

This is something I think about for AI engineering as well, which is the big labs want you to hand over everything in the prompts, and only code of English, and then the smaller brains, the GPU pours, always want to write more code to make things more deterministic and reliable and controllable.

“This is something I think about for AI engineering as well, which is the big labs want you to hand over everything in the prompts, and only code of English, and then the smaller brains, the GPU pours, always want to write more code to make things more deterministic and reliable and controllable.”
Speaker
Shawn Wang
Publisher
Latent Space

57 / evaluation

The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.

“The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.”
Speaker
Shawn Wang
Publisher
Latent Space

59 / evaluation

I will say I've been dealing with EDB a little bit from my conference, and they've been extremely responsive and it's been nice to see, because I never get to see this out of government, nice to see that as someone that wants to bring a foreign business into Singapore, they're kind of rolling on the welcome mat.

“I will say I've been dealing with EDB a little bit from my conference, and they've been extremely responsive and it's been nice to see, because I never get to see this out of government, nice to see that as someone that wants to bring a foreign business into Singapore, they're kind of rolling on the welcome mat.”
Speaker
Shawn Wang
Publisher
Latent Space

60 / evaluation

I'm pretty science-based, like, you know, but probably the most like spiritual woo-woo thing about me is I don't think that would lead to consciousness or AGI just because like there's something in- there's a soul, you know?

“I'm pretty science-based, like, you know, but probably the most like spiritual woo-woo thing about me is I don't think that would lead to consciousness or AGI just because like there's something in- there's a soul, you know?”
Speaker
Shawn Wang
Publisher
Latent Space

62 / evaluation

You're one of like, you're maybe the first PhD thesis defense I've ever watched in like this AI world, because most people just publish single papers, but every paper of yours is a banger.

“You're one of like, you're maybe the first PhD thesis defense I've ever watched in like this AI world, because most people just publish single papers, but every paper of yours is a banger.”
Speaker
Shawn Wang
Publisher
Latent Space

63 / evaluation

I will mostly agree and I'll slightly disagree in terms of this, which is like, whether designing for humans also overlaps with designing for AI. So Malte Ubo, who's the CTO of Vercel, who is creating basically JavaScript's competitor to LangChain, they're observing that basically, like if the API is easy to understand for humans, it's actually much easier to understand for LLMs, for example, because they're not overloaded functions.

“I will mostly agree and I'll slightly disagree in terms of this, which is like, whether designing for humans also overlaps with designing for AI. So Malte Ubo, who's the CTO of Vercel, who is creating basically JavaScript's competitor to LangChain, they're observing that basically, like if the API is easy to understand for humans, it's actually much easier to understand for LLMs, for example, because they're not overloaded functions.”
Speaker
Shawn Wang
Publisher
Latent Space

67 / evaluation

You've done very well. And I think you've honestly done the community a service by reading all these papers so that we don't have to, because the joke is often that, you know, what is one prompt is like then inflated into like a 10 page PDF that's posted on archive.

“You've done very well. And I think you've honestly done the community a service by reading all these papers so that we don't have to, because the joke is often that, you know, what is one prompt is like then inflated into like a 10 page PDF that's posted on archive.”
Speaker
Shawn Wang
Publisher
Latent Space

70 / evaluation

These are the same things to the model. That is a huge, huge win for interpretability, because up to now, we were only doing interpretability on toy models, like a few million parameters, a model of Go or chess or whatever.

“These are the same things to the model. That is a huge, huge win for interpretability, because up to now, we were only doing interpretability on toy models, like a few million parameters, a model of Go or chess or whatever.”
Speaker
Shawn Wang
Publisher
Latent Space

72 / evaluation

The other thing is RLHF came from the alignment community. And I think there's a lot of conception that maybe it's due to safety concerns, but I feel like it's really over the past two, three years expanded to just this produces a better model period, even if you don't really are not that concerned about existential risk.

“The other thing is RLHF came from the alignment community. And I think there's a lot of conception that maybe it's due to safety concerns, but I feel like it's really over the past two, three years expanded to just this produces a better model period, even if you don't really are not that concerned about existential risk.”
Speaker
Shawn Wang
Publisher
Latent Space

73 / evaluation

We're basically saying that, I think that one of the lessons from AlphaGo is that people thought that human interest in Go would be diminished because computers are better than humans.

“We're basically saying that, I think that one of the lessons from AlphaGo is that people thought that human interest in Go would be diminished because computers are better than humans.”
Speaker
Shawn Wang
Publisher
Latent Space

74 / evaluation

I don't know how to describe, you've done so much work in a very short amount of time at Meta, but you were most notably leading Llama 2 and now today we're also coordinating on the release of Llama 3.

“I don't know how to describe, you've done so much work in a very short amount of time at Meta, but you were most notably leading Llama 2 and now today we're also coordinating on the release of Llama 3.”
Speaker
Shawn Wang
Publisher
Latent Space

75 / evaluation

Because I didn't know that, I don't know how much to believe, you know, like there's a lot of these kinds of papers where it makes a lot of noise, but it doesn't actually pan out.

“Because I didn't know that, I don't know how much to believe, you know, like there's a lot of these kinds of papers where it makes a lot of noise, but it doesn't actually pan out.”
Speaker
Shawn Wang
Publisher
Latent Space

79 / evaluation

Yeah, so I really like this concept of evaluation. So actually, yeah, I think there's typically what I always say is like sort of 25 is random chance, 50 is average human, 75 is expert human, 90 is you're cheating.

“Yeah, so I really like this concept of evaluation. So actually, yeah, I think there's typically what I always say is like sort of 25 is random chance, 50 is average human, 75 is expert human, 90 is you're cheating.”
Speaker
Shawn Wang
Publisher
Latent Space

80 / evaluation

Obviously, I think you're like our second or third person from Hugging Face on the podcast and it's like the definitional sort of open AI company, maybe the real open AI.

“Obviously, I think you're like our second or third person from Hugging Face on the podcast and it's like the definitional sort of open AI company, maybe the real open AI.”
Speaker
Shawn Wang
Publisher
Latent Space

82 / evaluation

Like, or So it's a very interesting observation where like, most efficiency work is just busy work, or like, it's work at a small scale that doesn't, that just ignores the fact that like, this thing doesn't scale, because you haven't scaled it.

“Like, or So it's a very interesting observation where like, most efficiency work is just busy work, or like, it's work at a small scale that doesn't, that just ignores the fact that like, this thing doesn't scale, because you haven't scaled it.”
Speaker
Shawn Wang
Publisher
Latent Space

84 / evaluation

One thing I'll, one thing I'll mention quickly is that a lot of the stuff that you mentioned is typically not part of the normal interview loop. It's actually really hard to interview for because this is the stuff that you polish out in, as you go into production, the coding interviews are typically about the happy path.

“One thing I'll, one thing I'll mention quickly is that a lot of the stuff that you mentioned is typically not part of the normal interview loop. It's actually really hard to interview for because this is the stuff that you polish out in, as you go into production, the coding interviews are typically about the happy path.”
Speaker
Shawn Wang
Publisher
Latent Space

85 / evaluation

I have some appreciation, I think when you had me on your podcast, I was still working at Temporal and that was like a nice Framework, if you live within Temporal's boundaries, you can pretend that all those faults don't exist, and you can, you can code in a sort of very fault tolerant way.

“I have some appreciation, I think when you had me on your podcast, I was still working at Temporal and that was like a nice Framework, if you live within Temporal's boundaries, you can pretend that all those faults don't exist, and you can, you can code in a sort of very fault tolerant way.”
Speaker
Shawn Wang
Publisher
Latent Space

86 / evaluation

If some AI engineer would not know, I don't know what, , I don't know where we would stoop to, to call something required knowledge, , or you're not part of the cool kids club.

“If some AI engineer would not know, I don't know what, , I don't know where we would stoop to, to call something required knowledge, , or you're not part of the cool kids club.”
Speaker
Shawn Wang
Publisher
Latent Space

88 / evaluation

I, I do often say that I think AI engineering is about 90 percent software engineering with like the, the 10 percent of like really strong really differentiated AI engineering.

“I, I do often say that I think AI engineering is about 90 percent software engineering with like the, the 10 percent of like really strong really differentiated AI engineering.”
Speaker
Shawn Wang
Publisher
Latent Space

89 / evaluation

One thing I see in the recent papers that have been coming out is this sort of concept of multi-stage training data. And if you're doing full fine tuning, maybe the move or the answer is don't train 500 billion tokens on just code, because then yeah, it's going to massively overfit to just code.

“One thing I see in the recent papers that have been coming out is this sort of concept of multi-stage training data. And if you're doing full fine tuning, maybe the move or the answer is don't train 500 billion tokens on just code, because then yeah, it's going to massively overfit to just code.”
Speaker
Shawn Wang
Publisher
Latent Space

90 / evaluation

Yeah, I think, you know, the one thing that makes this sort of generative AI era very different from the sort of data science-y type era is that it is very non-deterministic and it's hard to control.

“Yeah, I think, you know, the one thing that makes this sort of generative AI era very different from the sort of data science-y type era is that it is very non-deterministic and it's hard to control.”
Speaker
Shawn Wang
Publisher
Latent Space

91 / evaluation

It's just retrieval. And here it's like, The home of generative AI, this, whatever hyperstition is in my mind, like this is actually pushing the edge of what generative and creativity in AI means.

“It's just retrieval. And here it's like, The home of generative AI, this, whatever hyperstition is in my mind, like this is actually pushing the edge of what generative and creativity in AI means.”
Speaker
Shawn Wang
Publisher
Latent Space

92 / evaluation

When you say you need a certain set of tools for people to sort of invent things from first principles Devin is the agent that I think has been able to utilize its tools very effectively.

“When you say you need a certain set of tools for people to sort of invent things from first principles Devin is the agent that I think has been able to utilize its tools very effectively.”
Speaker
Shawn Wang
Publisher
Latent Space

95 / evaluation

Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.

“Basically, I think I'm very just impressed by how first principles, your ideas around what the workflow is. And I think that's why you're not as reliant on like the LLM improving, because it's actually just about improving the workflow that you would recommend to people.”
Speaker
Shawn Wang
Publisher
Latent Space
Search evidence