High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

John Schulman: evaluation

15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

“I would say there's a decent amount of room for variation in exactly how you do the training process.”

— John Schulman

Source trail

Everything needed to verify it.

Speaker
John Schulman
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 May 2024
Publisher
Dwarkesh Podcast

Transcript context

…A couple of rapid-fire questions about RLHF. Obviously, RLHF is important to make these models useful. So maybe the "lobotomized" description is inaccurate. However, there is a sense in which all of these models, once they're put in a chatbot form, have a very similar way of speaking. They really want to “delve” into things. They want to turn things into bullet points. They often seem to have this formal and dull way of speaking. There are complaints that they're not as creative. Like we were talking about before, they could only do rhyming poetry and not non-rhyming poetry until recently. Is that a result of the particular way in which RLHF happens now? If so, is it because of who the raters are? Is it because of what the loss function is? Why is this the way all chatbots look? I would say there's a decent amount of room for variation in exactly how you do the training process. We're actively trying to improve this and make the writing more lively and fun. We've made some progress like improving the personality of ChatGPT. It is more fun and it's better when you're trying to chit chat with it and so forth. It's less robotic. It's an interesting question how some of the ticks came about, like the word "delve." I've actually caught myself using that word recently. I don't know if it rubbed off on me from the model. Actually, there might also be some funny effects going on where there's unintentional distillation happening between the language model and providers. If you hire someone to go do a labeling task, they might just be feeding it into a model. They might be pulling up their favorite chatbot, feeding it in, having the model do the task, and then copying and pasting it back. So that might account for some of the convergence. Some of the things we're seeing are just what people like. People do like bullet points. They like structured responses. People do often like the big info dumps that they get from the models. So it's not completely clear how much is just a quirk of the particular choices and design of the post-training processes, and how much is actually intrinsic to what people actually want. It does seem persistently more verbose than some people want. Maybe it’s just because during the labeling stage, the raters will prefer the more verbose answer. I wonder if it's inherent because of how it's pre-trained and the stop sequence doesn't come up that often and it really wants to just keep going.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence