Speakers in the public record
Claim mix
belief 12evaluation 10uncertainty 4prediction 3commitment 1observation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
31 published records
“Now, the last thing I want to add is having bigger models enables us to collect better data, for instance, at LHF stage, because that's the model we use for the annotation.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think we can maybe transition towards some of your personal stuff. We kept you here for a long time.”
- Publisher
- Latent Space
“I would say the diffusion people have actually started to swing back to pixel level and probably that will presage the language people also moving towards, you know, 1 million vocabulary and then, you know, whatever the natural limit is for character level.”
- Publisher
- Latent Space
“If you think about the evolution of the models, I think up until Llama3, with Meta AI and some of these things, I'm like, it makes sense that they want to build their own models and they're multi-modal.”
- Publisher
- Latent Space
“The other thing is, I think a dense model is just one specific variation of the model for an hyperparameter for anMoEwith basically one expert.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“We still need to have, at the architecture level, some kind of variable inference length thing that lets you actually think in latent space, like you're talking about. I don't know if there's any papers that you're thinking about.”
- Publisher
- Latent Space
“I think we can do a lot better in the future and not just like with transformers, but for instance, to me, like it doesn't make sense to use the same amount of compute per token for every token.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“For instance, if your model can calculate something it was wrong before and now it has access to a calculator and you can retrain your model on that, then you're learning something new.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I mean, what version, but by far compared to the version originally released, even now, I think there's maybe the last clouds on a 3.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think folks were asking if you see that as an interesting direction to kind of having specific synthetic data generation things.”
- Publisher
- Latent Space
“Improving on helpfulness, which is one of the main dimensions that people look at, I think, in the arena, which is, by the way, a very interesting evaluation. Because when we did the preview, and I don't know yet what will be the results for this new Llama 3, but we ended very high in this blind test leaderboard.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“To be fair, I think OpenAI knew that at the time of Chinchilla paper, but yeah, basically Chinchilla said we have to revisit the scaling laws originally published by Kepler and emphasize much more the importance of training tokens.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“We are state-of-the-art there. I think the model will be pretty good at that. We have a lot of gems about tools in the paper, but the model is fine-tuned to do tool usage, to zero-shot function calling.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“One of our previous guests, Brian Bischoff, is also asking about like, how do we think about evals for practical things like confidence estimation, structured output, you know, stuff like that.”
- Publisher
- Latent Space
“If I'm pretty sure that if I call them like, RLHF without human in the loop, but like a discriminator which is synthetic human in the loop, I will have get much more citations today.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“To keep collecting new data and new tokens, people are saying we are lacking of tokens, but if you think about those kinds of tokens, where the model always goes to correct its own weakness, it can say, that's 10 plus 10, that's an easy example, probably the model knows, but imagine for something more complex, 10 plus 10, I expect this to be 20.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think in the future, it should hopefully go more as this is a task and I return it.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“Overall, like scaling, I don't know if it's all you need, but I will not bet against scaling.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“The intuition in general is like, for instance, for code, because this is factual, you can check if the code is correct or not, RLHF is not the way to go.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I don't know about this model exactly, but I think like LlamaT had better performance overall.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think there's a limit with respect that we grow with respect to the model size.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think there's a lot of chatter obviously about synthetic data and like there was the Rephrase the Web paper that came out maybe a few months ago about using, you know, Mastral to make training data better.”
- Publisher
- Latent Space
“We love just Meta's commitment to open source and, you know, you do what you need to do to make it work for your organization.”
- Publisher
- Latent Space
“The other thing is RLHF came from the alignment community. And I think there's a lot of conception that maybe it's due to safety concerns, but I feel like it's really over the past two, three years expanded to just this produces a better model period, even if you don't really are not that concerned about existential risk.”
- Publisher
- Latent Space
“To me, you know, I'm not exactly sure what you guys did, but like, I feel like when people say synthetic data, there needs to be different categories of synthetic data now, because I think there's so many different usage of this thing.”
- Publisher
- Latent Space
“We're basically saying that, I think that one of the lessons from AlphaGo is that people thought that human interest in Go would be diminished because computers are better than humans.”
- Publisher
- Latent Space
“I don't know how to describe, you've done so much work in a very short amount of time at Meta, but you were most notably leading Llama 2 and now today we're also coordinating on the release of Llama 3.”
- Publisher
- Latent Space
“Because I didn't know that, I don't know how much to believe, you know, like there's a lot of these kinds of papers where it makes a lot of noise, but it doesn't actually pan out.”
- Publisher
- Latent Space
“I think there was obviously the Kepler, and then there was Chinchilla, and then people kind of got the Llama scaling law, like the 100 to 200x parameter to token ratio.”
- Publisher
- Latent Space
“I think, yeah, very underrated, very underrated, this sort of PhD with industry expertise, because you're also publishing papers the whole time.”
- Publisher
- Latent Space
“Well, I think your progress into NLP was like really strong, because like the first thing you worked on at Meta was Bloom.”
- Publisher
- Latent Space