High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Michael Royzen

Published podcast speaker

Claims
15
Episodes
1
Shows
1
Named items
3

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

tool / uses

BERT

“I used the standard BERT and also Longformer, which came out around the same time. And Longformer was interesting because it had a much bigger context window than those models at the time, like BERT, all of the first gen encoder only models, they only had a context window of 512 tokens and it's fixed. There's none of this alibi or ROPE that we have now where we can basically massage it to be longer. They're fixed, 512 absolute encodings. Longformer at the time was the only way that you can fit, say, like a sequence length or ask a question about like 4,000 tokens worth of text.”

Latent Space · 3 Nov 2023

Evidence receipt · Source ↗

tool / uses

Longformer

“I used the standard BERT and also Longformer, which came out around the same time. And Longformer was interesting because it had a much bigger context window than those models at the time, like BERT, all of the first gen encoder only models, they only had a context window of 512 tokens and it's fixed. There's none of this alibi or ROPE that we have now where we can basically massage it to be longer. They're fixed, 512 absolute encodings. Longformer at the time was the only way that you can fit, say, like a sequence length or ask a question about like 4,000 tokens worth of text.”

Latent Space · 3 Nov 2023

Evidence receipt · Source ↗

other / likes

Nvidia

“And we start working with Nvidia, which is great. And something that I love about Nvidia, by the way, is that after that intro, we got matched with like a dedicated team.”

Latent Space · 3 Nov 2023

Evidence receipt · Source ↗

Claim ledger

What Michael said.

15 transcript-backed records

01 / belief

Yeah, sorry. So we filtered Common Crawl just by the top, I think, 10,000, just to limit this, because obviously there's this massive long tail of small sites that are really cool, actually.

“Yeah, sorry. So we filtered Common Crawl just by the top, I think, 10,000, just to limit this, because obviously there's this massive long tail of small sites that are really cool, actually.”
Speaker
Michael Royzen
Publisher
Latent Space

04 / belief

My take is that it's much easier having been on both sides of that coin now, it's much easier to stay obsessed every single day when the genesis of your startup is something that really spoke to you in an incredibly meaningful way beyond just being some insight that you've noticed.

“My take is that it's much easier having been on both sides of that coin now, it's much easier to stay obsessed every single day when the genesis of your startup is something that really spoke to you in an incredibly meaningful way beyond just being some insight that you've noticed.”
Speaker
Michael Royzen
Publisher
Latent Space

09 / preference

And we start working with Nvidia, which is great. And something that I love about Nvidia, by the way, is that after that intro, we got matched with like a dedicated team.

“And we start working with Nvidia, which is great. And something that I love about Nvidia, by the way, is that after that intro, we got matched with like a dedicated team.”
Speaker
Michael Royzen
Publisher
Latent Space

10 / evaluation

In fact, it's a little challenging sometimes to like finish kind of like the rest of like the description of your pitch because like, he'll just like asking all these questions about how it works.

“In fact, it's a little challenging sometimes to like finish kind of like the rest of like the description of your pitch because like, he'll just like asking all these questions about how it works.”
Speaker
Michael Royzen
Publisher
Latent Space

11 / evaluation

Even GPT, like the text DaVinci 2 was available at the time, wasn't that good at generating code and it would generate like very, very short, very incomplete code snippets. And so we launched that last summer, got some traction, but really like we were only doing like, I don't know, maybe like 10,000 searches a day.

“Even GPT, like the text DaVinci 2 was available at the time, wasn't that good at generating code and it would generate like very, very short, very incomplete code snippets. And so we launched that last summer, got some traction, but really like we were only doing like, I don't know, maybe like 10,000 searches a day.”
Speaker
Michael Royzen
Publisher
Latent Space

12 / commitment

That's what we've been doing. So we launched the very first version of Find in its current incarnation after like the previous demo connected to our own index.

“That's what we've been doing. So we launched the very first version of Find in its current incarnation after like the previous demo connected to our own index.”
Speaker
Michael Royzen
Publisher
Latent Space

13 / evaluation

When I returned to it in fall of 2021, when BigScience released T0, when BigScience released the T0 models, that was a massive jump in the reasoning ability of the model. And it was better at reasoning, it was better at summarization, it was still a glorified summarizer basically.

“When I returned to it in fall of 2021, when BigScience released T0, when BigScience released the T0 models, that was a massive jump in the reasoning ability of the model. And it was better at reasoning, it was better at summarization, it was still a glorified summarizer basically.”
Speaker
Michael Royzen
Publisher
Latent Space

15 / recommendation

Mentions personal use of BERT. Mentions personal use of Longformer.

“I used the standard BERT and also Longformer, which came out around the same time. And Longformer was interesting because it had a much bigger context window than those models at the time, like BERT, all of the first gen encoder only models, they only had a context window of 512 tokens and it's fixed. There's none of this alibi or ROPE that we have now where we can basically massage it to be longer. They're fixed, 512 absolute encodings. Longformer at the time was the only way that you can fit, say, like a sequence length or ask a question about like 4,000 tokens worth of text.”
Speaker
Michael Royzen
Publisher
Latent Space
Search evidence