Evidence receipt / belief
Published · transcript-backedNathan Lambert: belief
11 Jan 2024 Latent Space RLHF 201 - with Nathan Lambert of AI2 and Interconnects
“I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Lambert
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Jan 2024
- Publisher
- Latent Space
Transcript context
…Does this change as you think about a chain of thought reasoning and things like that? Like does the state become part of the chain that you're going through? There's work that I've mentioned on one slide called process reward models that essentially rewards each step in the chain of thought reasoning. It doesn't really give the part of interaction, but it does make it a little bit more fine grained where you can think about like calling it at least you have many states from your initial state. That formulation I don't think people have fully settled on. I think there's a bunch of great work out there, like even OpenAI is releasing a lot of this and let's verify step by step is there pretty great paper on the matter. I think in the next year that'll probably get made more concrete by the community on like if you can easily draw out like if chain of thought reasoning is more like RL, we can talk about that more later. That's a kind of a more advanced topic than we probably should spend all the time on. RLHF for decision making. You have a slide here that compares pre-deep RL versus deep RL.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.