High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Mark Huang

Published podcast speaker

Claims
14
Episodes
1
Shows
1
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

hardware / uses

GPT-4

“We used GPT-4 to rephrase certain aspects of the chat data, reformatting it or kind of generating new types of tokens and language and types of data that the model could see.”

Latent Space · 30 May 2024

Evidence receipt · Source ↗

Claim ledger

What Mark said.

6 transcript-backed records

02 / evaluation

Underrated specific instance would be the DeepSeek paper where I'd never seen it before, but the multi-head latent attention. That was really unexpected to me because I thought I'd seen every way that people wanted to cut mixture of experts into interesting ways.

“Underrated specific instance would be the DeepSeek paper where I'd never seen it before, but the multi-head latent attention. That was really unexpected to me because I thought I'd seen every way that people wanted to cut mixture of experts into interesting ways.”
Speaker
Mark Huang
Publisher
Latent Space

03 / evaluation

Obviously, you know, you don't build everything that people ask you to build, but we know what's useful, right? Because I think that you're totally right there.

“Obviously, you know, you don't build everything that people ask you to build, but we know what's useful, right? Because I think that you're totally right there.”
Speaker
Mark Huang
Publisher
Latent Space

04 / evaluation

Yeah, I think there's a huge resurgence in what I would call model alchemy to a certain extent, because you're taking all of these LoRa's and you're mixing them together.

“Yeah, I think there's a huge resurgence in what I would call model alchemy to a certain extent, because you're taking all of these LoRa's and you're mixing them together.”
Speaker
Mark Huang
Publisher
Latent Space

05 / evaluation

Yeah, in terms of, you know, all the literature out there, I would say, honestly, it's probably still TBD as to like the trade offs between the approach we did, which is more of a curriculum learning approach after the fact versus inherently training a model with a long context throughout, because I just don't think people have looked at the scaling properties of it in deep detail.

“Yeah, in terms of, you know, all the literature out there, I would say, honestly, it's probably still TBD as to like the trade offs between the approach we did, which is more of a curriculum learning approach after the fact versus inherently training a model with a long context throughout, because I just don't think people have looked at the scaling properties of it in deep detail.”
Speaker
Mark Huang
Publisher
Latent Space

06 / evaluation

We do have historical precedent, where the original code bomb was trained further from Mama 2, and it just lost all its language capability, basically, right? So I don't want to call that project like deem it as a failure, but it wasn't a really successful generalization exercise, because, you know, these models are about flexibility and being like generic to a certain extent.

“We do have historical precedent, where the original code bomb was trained further from Mama 2, and it just lost all its language capability, basically, right? So I don't want to call that project like deem it as a failure, but it wasn't a really successful generalization exercise, because, you know, these models are about flexibility and being like generic to a certain extent.”
Speaker
Mark Huang
Publisher
Latent Space
Search evidence