Speakers in the public record
Claim mix
belief 16evaluation 8uncertainty 5prediction 1commitment 1recommendation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
32 published records
“As reinforcement learning is so much less compute, like it is a richer signal in terms of its impact. Because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner.”
- Publisher
- Latent Space
“Anyway, we're here to talk about RLHF 101. You did a presentation, and I think you expressed some desire to rerecord it.”
- Publisher
- Latent Space
“Like I know some that are in the like the RLHF as a service space will become busy. I think for good reason, just because like.”
- Publisher
- Latent Space
“I think if you zoom into any of the details to look at like the agreement number, so how if you look at a test set, you'll have a chosen and rejected and you can take the reward model you're training, pass in those completions and you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward.”
- Publisher
- Latent Space
“Oh, I don't even know what inst means, but just saying like they use their adjective that they like. I think Entropic also like steerable is another one.”
- Publisher
- Latent Space
“I think big labs are indexed on their own base models so they don't know like what's swapping between CloudBase or GPT-4 base how that would change any notion of preference or what you do with RLHF.”
- Publisher
- Latent Space
“I don't know if this is a relevant comparison for you, but OpenAI also recently released a weak to strong generalization paper where they actually talked about a few intermediate checkpoints for GPT-4.”
- Publisher
- Latent Space
“I think the things that people see now is like the small models don't really handle nuance as well and they could be more repetitive even if they have really good instruction tuning.”
- Publisher
- Latent Space
“I think the main result that I think most people talk about at this stage, we're talking about September 2020 and then going into, I guess maybe last year was Vicuña as one of the more interesting applications of instruction tuning that pushed LLAMA1 from, let's say a GPT 3-ish model to a GPT 3.”
- Publisher
- Latent Space
“I think their papers are sometimes pretty funny because they're not capabilities papers.”
- Publisher
- Latent Space
“People often say, I don't know what I want, but I'll know when I see it. This is that expressed in reinforcement learning tools.”
- Publisher
- Latent Space
“I think with the language model, it's very hard to define what an environment is.”
- Publisher
- Latent Space
“I think the reason why it's not really talked about is just because the RLHF techniques that people use were built in labs like OpenAI and DeepMind where there are some of these people.”
- Publisher
- Latent Space
“I think people want to hear it. I think there's a lot of higher level explanations out there.”
- Publisher
- Latent Space
“I try to do it every six or 12 months is my estimated cadence, just to refine the ways that I say things. And people will see that we don't know that much more, but we have a bit of better way of saying what we don't know.”
- Publisher
- Latent Space
“I would say like it does remind me of FlashAttention a little bit in a sense that like kind of an equivalent thing to the thing it's replacing and it's just faster, cheaper, just better.”
- Publisher
- Latent Space
“Can we, since you mentioned expensiveness, I think you may have joined one of our spaces back in Lama 2 was released.”
- Publisher
- Latent Space
“People release checkpoints, but that's how we should be thinking about it because the optimizer is so strong and it's like we don't know what's happening in this kind of intermediate land.”
- Publisher
- Latent Space
“I think if the people are kind of locked into using synthetic data, people also think that synthetic data is like GPT-4 is more accurate than humans at labeling preferences.”
- Publisher
- Latent Space
“We can dive right in. I don't know if there's any other topics that we want to lay out as groundwork.”
- Publisher
- Latent Space
“I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later.”
- Publisher
- Latent Space
“I think I saw people criticizing it for like just being like safety washing from the fact that they're like talking about GPT-2 still, which is such a kind of like odd model to focus on.”
- Publisher
- Latent Space
“It's really tricky to actually do that. I think that people just keep using GPT-4 because it's really cheap.”
- Publisher
- Latent Space
“I think InstructGPT does something where they like try to get the RL model to match the instruction tuning dataset because they were really happy with that dataset to constrain the distribution.”
- Publisher
- Latent Space
“I think the reason this really is done on a deep level is that you're not actually trying to model any contestable preference in this.”
- Publisher
- Latent Space
“I think in the long run, it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about.”
- Publisher
- Latent Space
“Is this text bad? That's not that surprising, I think, because you could use like a hundred times smaller language model and do much better at filtering than RLHF.”
- Publisher
- Latent Space
“However, reinforcement learning proved highly effective, particularly given its cost and time effectiveness. So you don't really know exactly what the costs and time that Meta is looking at, because they have a huge team and a pretty good amount of money here to release these Llama models.”
- Publisher
- Latent Space
“There's an early one on RLHF, which is, this stuff is all just like when I figure it out in my brain. So I wrote an article that's like how RLHF actually works, which is just the intuitions that I had put together in the summer about RLHF, and that was pretty well.”
- Publisher
- Latent Space
“Like, I don't know exactly if they're doing this now, but you can kind of see why doing RLHF at scale and prioritizing a lot of different endpoints would be hard because these are all things I'd be interested in if I was scaling up a big team to do RLHF and like what is going into the preference data.”
- Publisher
- Latent Space
“There's essentially the thing in my mind that I can't get past is the difference between the control you get in training a reward model and then training a policy because essentially everything you want your reward model to do might not be everything that you train the policy to do in the RLHF step where you have like the two different prompt distributions.”
- Publisher
- Latent Space
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche, because you're not going to make your model chat GPT like better than OpenAI or anything like that.”
- Publisher
- Latent Space