Evidence receipt / commitment
Published · transcript-backedSteve Hsu: commitment
23 Aug 2022 Dwarkesh Podcast Steve Hsu - Intelligence, Embryo Selection, & The Future of Humanity
“The very first paper we wrote on this was to prove that using accurate genomic data and these very abstract theorems in combination could predict how much data you need to “solve” individual human traits.”
Source trail
Everything needed to verify it.
- Speaker
- Steve Hsu
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 23 Aug 2022
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, so the super fascinating thing about this is that the diseases you talked about—or at least their risk profiles—are polygenic. You can have thousands of SNPs (single nucleotide polymorphisms) determining whether you will get a disease. So, I'm curious to learn how you were able to transition to this space and how your knowledge of mathematics and physics was able to help you figure out how to make sense of all this data. Yeah, that's a great question. So again, I was stressing the fundamental scientific importance of all this stuff. If you go into a slightly higher level of detail—which you were getting at with the individual SNPs, or polymorphisms—there are individual locations in the genome, where I might differ from you, and you might differ from another person. Typically, each pair of individuals will differ at a few million places in the genome—and that controls why I look a little different than you A lot of times, theoretical physicists have a little spare energy and they get tired of thinking about quarks or something. They want to maybe dabble in biology, or they want to dabble in computer science, or some other field. As theoretical physicists, we always feel, “Oh, I have a lot of horsepower, I can figure a lot out.” (For example, Feynman helped design the first parallel processors for thinking machines.) I have to figure out which problems I can make an impact on because I can waste a lot of time. Some people spend their whole lives studying one problem, one molecule or something, or one biological system. I don't have time for that, I'm just going to jump in and jump out. I'm a physicist. That's a typical attitude among theoretical physicists. So, I had to confront sequencing costs about ten years ago because I knew the rate at which they were going down. I could anticipate that we’d get to the day (today) when millions of genomes with good phenotype data became available for analysis. A typical training run might involve almost a million genomes or half a million genomes. The mathematical question then was: What is the most effective algorithm given a set of genomes and phenotype information to build the best predictor? This can be boiled down to a very well-defined machine learning problem. It turns out, for some subset of algorithms, there are theorems— performance guarantees that give you a bound on how much data you need to capture almost all of the variation in the features. I spent a fair amount of time, probably a year or two, studying these very famous results, some of which were proved by a guy named Terence Tao, a Fields medalist. These are results on something called compressed sensing: a penalized form of high dimensional regression that tries to build sparse predictors. Machine learning people might notice L1-penalized optimization. The very first paper we wrote on this was to prove that using accurate genomic data and these very abstract theorems in combination could predict how much data you need to “solve” individual human traits. We showed that you would need at least a few hundred thousand individuals and their genomes and their heights to solve for height as a phenotype. We proved that in a paper using all this fancy math in 2012. Then around 2017, when we got a hold of half a million genomes, we were able to implement it in practical terms and show that our mathematical result from some years ago was correct. ncy math in 2012. Then around 2017, when we got a hold of half a million genomes, we were able to implement it in practical terms and show that our mathematical result from some years ago was correct. The transition from the low performance of the predictor to high performance (which is what we call a “phase transition boundary” between those two domains) occurred just where we said it was going to occur. Some of these technical details are not understood even by practitioners in computational genomics who are not quite mathematical. They don't understand these results in our earlier papers and don't know why we can do stuff that other people can't, or why we can predict how much data we'll need to do stuff. It's not well-appreciated, even in the field. But when the big AI in our future in the singularity looks back and says, “Hey, who gets the most credit for this genomics revolution that happened in the early 21st century?” they're going to find these papers on the archive where we proved this was possible, and how five years later, we actually did it. Right now, it's under-appreciated, but the future AI––that Roko's Basilisk AI–will look back and will give me a little credit for it.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.