Evidence receipt / prediction
Published · transcript-backedJohn Schulman: prediction
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“I realized a lot of the things that people thought were flaws in language models, like blatant hallucination, could be not completely fixed but things that you could make a lot of progress on with pretty straightforward methods.”
Source trail
Everything needed to verify it.
- Speaker
- John Schulman
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Before ChatGPT, OpenAI had these instruction following models. The idea there was that we had base models that people could prompt them in elaborate ways. But they were also hard to prompt. They basically do autocomplete so you had to set up a very good prompt with some examples. People at OpenAI were working on just taking the base models and making them easier to prompt. So if you just wrote a question it would answer the question, instead of giving you more questions or something. So we had these instruction following models, which were like base models but a little easier to use. Those are the original ones deployed in the API. Or after GPT-3, those were the next generation of models. At the same time there were definitely a lot of people thinking about chat. Google had some papers like LaMDA and earlier, Meena. They had these chatbots. It was more like a base model that was really specialized to the task of chat. It was really good at chat. Looking at the examples from the paper, it was more used for fun applications where the model would take on some persona and pretend to be that persona. It was not so functional where it could help me refactor my code. So there were definitely people thinking about chat. I had worked before on a project looking at chat called WebGPT, which was more about doing question answering with the help of web browsing and retrieval. When you do question answering, it really wants to be in a chat. You always want to ask follow-up questions or sometimes the model should ask a clarifying question because the question is ambiguous. It was clear after we did the first version, that the next version should be conversational. So we started working on the conversational chat assistant. This was built on top of GPT-3.5, which was done training at the beginning of 2022. That model was quite good at language and code. We quickly realized that it was actually quite good at coding help. That was one of the things we were excited about. We worked on that for most of the year. We had browsing as another feature in it although we ended up deemphasizing that later on because the model's internal knowledge was so good. The browsing wasn't the most interesting thing about it. We had it out to friends and family for a while and we were thinking about doing a public release. Actually, GPT-4 finished training in August that year. The flagship RL effort at OpenAI was the instruction following effort because those were the models that were being deployed into production. The first fine-tunes of GPT-4 used that whole stack. Those models were really good and everyone got really excited about that after seeing the instruct fine tune GPT-4s. They were really good. They would occasionally give you amazing outputs, but the model was clearly also pretty unreliable. It would sometimes hallucinate it a lot. It would sometimes give you pretty unhinged outputs. So it was clearly not quite ready for prime time, but it was obviously very good. also pretty unreliable. It would sometimes hallucinate it a lot. It would sometimes give you pretty unhinged outputs. So it was clearly not quite ready for prime time, but it was obviously very good. People forgot about chat for a little while after that, this alternative branch. We pushed it further and we ended up mixing together all the datasets, the instruct and the chat data, to try to get something that was the best of both worlds. The chat models were clearly easier to use. It automatically had much more sensible behavior in terms of the model knowing its own limitations. That was actually one of the things that I got excited about as we were developing it. I realized a lot of the things that people thought were flaws in language models, like blatant hallucination, could be not completely fixed but things that you could make a lot of progress on with pretty straightforward methods. The other thing about chat was when we had these instruct models. The task of “complete this text, but in a nice or helpful way” is a pretty poorly defined task. That task is both confusing for the model and for the human who's supposed to do the data labeling. Whereas for chat, people had an intuitive sense of what a helpful robot should be like. So it was just much easier for people to get an idea of what the model was supposed to do. As a result, the model had a much more coherent personality and it was much easier to get pretty sensible behavior robustly. Interesting. Is it the case that anybody could have made ChatGPT using your publicly available fine-tuning API?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.