High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

23 Jul 2024 Latent Space Llama 2, 3 & 4: Synthetic Data, RLHF, Agents on the path to Open Source AGI

“I think there's a limit with respect that we grow with respect to the model size.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
23 Jul 2024
Publisher
Latent Space

Transcript context

…Now that you're saying this, right, okay, so now we have 100k vocab, 200k vocab. Do we see a million vocab? Do we see infinity, which is no tokenizer, you know, like what's the natural limit of tokenization? Yeah. That's a good question. I don't know. I think there's a limit with respect that we grow with respect to the model size. So bigger models means possibly bigger vocabulary without affecting too much the training. But yeah, there's a lot of people, that's not my domain of expertise, but a lot of people are discussing the interest of having this kind of tokenizer, which doesn't fit like natural. Could we go to character level tokenizer? Could we go to actually multimodal tokenizer, which will like decompose at pixel level? I don't know. Future directions that could be very promising. I would say the diffusion people have actually started to swing back to pixel level and probably that will presage the language people also moving towards, you know, 1 million vocabulary and then, you know, whatever the natural limit is for character level.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence