Evidence receipt / uncertainty
Published · transcript-backedSteve Hsu: uncertainty
23 Aug 2022 Dwarkesh Podcast Steve Hsu - Intelligence, Embryo Selection, & The Future of Humanity
“What happens is that if you train a predictor in this ancestry group and go to a more distant ancestry group, there's a fall-off in the prediction quality. Again, this is a frontier question, so we don't know the answer for sure.”
Source trail
Everything needed to verify it.
- Speaker
- Steve Hsu
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 23 Aug 2022
- Publisher
- Dwarkesh Podcast
Transcript context
…Oh, okay. Interesting. That raises the next question I was about to ask: how applicable are these scores across different ancestral populations? Huge problem is that most of the data is from Europeans. What happens is that if you train a predictor in this ancestry group and go to a more distant ancestry group, there's a fall-off in the prediction quality. Again, this is a frontier question, so we don't know the answer for sure. But many people believe that there's a particular correlational structure in each population, where if I know the state of this SNP, I can predict the state of these neighboring SNPs. That is a product of that group's mating patterns and ancestry. Sometimes, the predictor, which is just using statistical power to figure things out, will grab one of these SNPs as a tag for the truly causal SNP in there. It doesn't know which one is genuinely causal, it is just grabbing a tag, but the tagging quality falls off if you go to another population (eg. This was a very good tag for the truly causal SNP in the British population. But it's not so good a tag in the South Asian population for the truly causal SNP, which we hypothesize is the same). It's the same underlying genetic architecture in these different ancestry groups. We don't know if that's a hypothesis. But even so, the tagging quality falls off. So my group spent a lot of our time looking at the performance of predictor training population A, and on distant population B, and modeling it trying to figure out trying to test hypotheses as to whether it's just the tagging decay that’s responsible for most of the faults. So all of this is an area of active investigation. It'll probably be solved in five years. The first big biobanks that are non-European are coming online. We're going to solve it in a number of years. Oh, what does the solution look like? Unless you can identify the causal mechanism by which each SNP is having an effect, how can you know that something is a tag or whether it's the actual underlying switch?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.