High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nat Friedman: evaluation

22 Mar 2023 Dwarkesh Podcast Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI

“I think Sam Altman did a good job of building that partnership because he knew that he needed access to the resources of a company like Microsoft to build large-scale AI and eventually AGI.”

— Nat Friedman

Source trail

Everything needed to verify it.

Speaker
Nat Friedman
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 Mar 2023
Publisher
Dwarkesh Podcast

Transcript context

…By the way, do you know why he knew that OpenAI would be worth investing at that point? I don't know. Actually, I've never asked him. That's a good question. I think OpenAI had already had some successes that were noticeable and I think, if your Satya and you're running this multi-trillion dollar company, you're trying to execute well and serve your customers but you're always looking for the next gigantic wave that is going to upend the technology industry. It's not just about trying to win cloud. It's – Okay, what comes after cloud? So you have to make some big bets and I think he thought AI could be one. And I think Kevin Scott deserves a lot of credit for really advocating for that aggressively. I think Sam Altman did a good job of building that partnership because he knew that he needed access to the resources of a company like Microsoft to build large-scale AI and eventually AGI. So I think it was some combination of those three people kind of coming together to make it happen. But I still think it was a very prescient bet. I've said that to people and they've said – Well, One billion dollars is not a lot for Microsoft. But there were a lot of other companies that could have spent a billion dollars to do that and did not. And so I still think that deserves a lot of credit. Okay, so GPT-3 comes out. I pinged Sam and Greg Brockman at OpenAI and they're like – Yeah, let's. We've already been experimenting with GPT-3 and derivative models and coding contacts. Let's definitely work on something. And to me, at least, and a few other people, it was not incredibly obvious what the product would be. Now, I think it's trivially obvious – Auto-complete, my gosh. Isn't that what the models do? But at the time my first thought was that it was probably going to be like a Q&A chatbot Stack Overflow type of thing. And so that was actually the first thing we prototyped. We grabbed a couple of engineers, SkyUga, who had come in from acquisition that we'd done, Alex Gravely, and started prototyping. The first prototype was a chatbot. What we discovered was that the demos were fabulous. Every AI product has a fantastic demo. You get this wow moment. It turns to maybe not be a sufficient condition for a product to be good. At the time the models were just not reliable enough, they were not good enough. I ask you a question 25% of the time you give me an incredible answer that I love. 75% of the time your answers are useless and are wrong. It's not a great product experience. And so then we started thinking about code synthesis. Our first attempts at this were actually large chunks of code synthesis, like synthesizing whole function bodies. And we built some tools to do that and put them in the editor. And that also was not really that satisfying. And so the next thing that we tried was to just do simple, small-scale auto-complete with the large models and we used the kind of IntelliSense drop down UI to do that. And that was better, definitely pretty good but the UI was not quite right. And we lost the ability to do this large scale synthesis. dels and we used the kind of IntelliSense drop down UI to do that. And that was better, definitely pretty good but the UI was not quite right. And we lost the ability to do this large scale synthesis. We still have that but the UI for that wasn't good. To get a function body synthesized you would hit a key. And then I don't know why this was the idea everyone had at the time, but several people had this idea that it should display multiple options for the function body. And the user would read them and pick the right one. And I think the idea was that we would use that human feedback to improve the model. But that turned out to be a bad experience because first you had to hit a key and explicitly request it. Then you had to wait for it. And then you had to read three different versions of a block of code. Reading one version of a block of code takes some cognitive effort. Doing it three times takes more cognitive effort. And then most often the result of that was like – None of them were good or you didn't know which one to pick. That was also like you're putting a lot of energy and you're not getting a lot out, sort of frustrating. Once we had that sort of single line completion working, I think Alex had the idea of saying we can use the cursor position in the AST to figure out heuristically whether you're at the beginning of a block and the code or not. And if it's not the beginning of a block, just complete a line. If it's the beginning of a block, show in line a full block completion. The number of tokens you request and when you stop gets altered automatically with no user interaction. And then the idea of using this sort of gray text like Gmail had done in the editor. So we got that implemented and it was really only kind of once all those pieces came together and we started using a model that was small enough to be low latency, but big enough to be accurate, that we reached the point where like the median new user loved Co-pilot and wouldn't stop using it. That took four months, five months, of just tinkering and sort of exploring. There were other dead ends that we had along the way. And then it became quite obvious that it was good because we had hundreds of internal users who were GitHub engineers. And I remember the first time I looked at the retention numbers, they were extremely high. It was like 60 plus percent after 30 days from first install. If you installed it, the chance that you were still using it after 30 is over 60 percent. And it's a very intrusive product. It's sort of always popping UI up and so if you don't like it, you will disable it. Indeed, 40 something percent of people did disable it but those are very high retention numbers for like an alpha first version of a product that you're using all day. Then I was just incredibly excited to launch it. And it's improved dramatically since then.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence