Temple and Browning model correlations of identity-by-descent (IBD) rates to derive analytical and simulation-based genome-wide significance thresholds for selection scans, apply these to TOPMed and UK Biobank cohorts, and show many signals cluster near structural-variant hotspots.
0:19Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Glad to be here for another deep dive.
0:30Yeah, so I want to start with something a bit abstract today. You know, whenever you step outside on a clear dark night and just look up. Oh, sure. Your brain immediately starts playing a trick on you.
0:43Like you don't just see a random scattering of stars. Right. You see like a hunter with a bow or a giant bear. Exactly. Or even just a perfectly straight line. We're deeply, I mean, fundamentally hardwired to find patterns, you know, to connect the dots in the dark.
0:59Yeah, absolutely. That ability to spot a pattern is, well, it's totally baked into our biology. Because historically, Kip is alive. So for sure. Recognizing a shape used to mean the difference between seeing a predator hidden in the brush or, you know, becoming its dinner.
1:13But in modern data science, that instinct can actually work against us. We desperately want to see a story in the noise. Which brings us to today's deep dive. We are taking that exact idea of connecting the dots and applying it to the ultimate map, which is the human genome.
1:28The biggest map we have. Right. And we are looking for the genetic constellations that tell the story of our recent evolution. We want to find those specific chunks of DNA that helped our ancestors survive.
1:40The literal biological signatures of the survival of the fittest. Which is a massive undertaking. It is. But what really happens when we look closely at the genome for signs of recent survival of the fittest, only to find our statistical math, might be tricking us into seeing evolutionary mirages.
1:56Yeah, that is the $1000000 question. Because when you feed massive data sets of human DNA into a computer, you inevitably find regions that look wildly unique. Just by random chance. Exactly. Certain genetic patterns will jump out and seem to suggest a profound evolutionary adaptation.
2:13But if your mathematical baseline for what counts as significant is flawed, you're basically just looking at a random cluster of stars and officially categorizing it as Orion. Wow. And, I mean, how could this change our understanding of human adaptation?
2:28Are the textbook examples of recent human evolution actually real, or are we just eager to connect random dots in the data? To untangle all of this? We really have to look at the very foundation of how we scan the genome for these adaptations in the 1st place.
2:42And we are leaning on some truly groundbreaking work to figure this out, right? We are. Today we celebrate the work of researchers from the University of Washington and the University of Michigan, specifically Seth D.
2:53Temple and Sharon R. Browning, who have advanced our understanding of statistical frameworks for genomic selection stands. Okay, so let's lay out what these researchers were actually looking for and, you know, why finding it is a monumental mathematical headache.
3:08Let's do it. When we talk about recent positive natural selection. We're essentially talking about a beneficial genetic trait that appears and then spreads through a population incredibly fast. Right. So consider a genetic mutation that suddenly gives a massive survival advantage.
3:23A classic example is lactase persistence. Oh, right, the ability to drink milk. Exactly. For most of human history. Your ability to digest milk basically turned off after childhood. But then, a mutation occurred allowing adults to digest lactose.
3:38It's massive in environments where dairy producing animals were available, especially during famine, being able to drink milk was an enormous caloric advantage. So the people carrying that mutation, survived, reproduced, and passed it on.
3:51Yeah. And in the grand timescale of human history, a trait like that sweeps through a population in the blink of an eye. And because it's spread so fast, it kind of drags the neighboring pieces of DNA along with it for the ride.
4:04Yes, exactly. That's the crucial part Every generation, human DNA is supposed to be shuffled. You get half your chromosomes from one parent, half from the other, and they cross over and break apart. The normal generational mixing.
4:17Right. But if a survival trait is racing through a population over just a few 100 generations, the surrounding genetic sequence just doesn't have time to be chopped up. So it stays intact. Exactly. This creates long, intact stretches of shared code that geneticists call identity by descent or IBD segments.
4:36Okay, so IBD segments are basically identical chunks of DNA shared between 2 completely different people today. Yeah, purely because they both inherited it from the same recent common ancestor who had that original survival mutation.
4:50Got it. So if a specific region of the genome is suddenly under strong positive selection, you'd see a massive spike in how often people share these long IBD segments in that exact spot. You've got it.
5:02So to find recent human evolution, you scan the entire genome, testing position after position, just looking for unusually high rates of these shared segments. But wait, if you are testing position after position across the entire genome.
5:15You aren't just running one test. You are running 10s of 1000s of statistical tests. Thousands and 1000s of them. And that introduces the multiple testing problem. If you run 50,000 separate statistical tests on any data set, random chance dictates that some of them are going to look highly significant.
5:31It's like rolling a 20 sided die. If you roll it once, getting a 20 is special. But if you roll it 50,000 times, you are going to see a mountain of 20s just by sheer volume. That is a perfect way to put it.
5:45If you don't correct your math to account for the sheer number of times you rolled the die, you will end up publishing a paper claiming you found a brilliant new evolutionary adaptation. When really, you just found statistical noise.
5:57Exactly. You just found a clump of 20s. Okay, let's unpack this. If these IBD segments are just shared chunks of DNA. Why is the math for figuring out statistical significance so much harder than standard genetic association studies?
6:10It all comes down to correlation. In a standard genome wide association study, or geewas, researchers look for links between genes and diseases. Right, and they fix the multiple testing problem by setting a notoriously strict P value threshold, like, what is it, 5 times 10 to the negative eight.
6:26Yeah, very strict. But in a GG Wells S, the genetic markers being tested are often treated as independent of each other. IBD segments by their very nature are long, overlapping physical stretches of chromosomes.
6:38Oh I see. Yeah, if you test a specific location on the chromosome for an IVD spike, and then you test another location, just microscopic fraction of a millimeter down the line, the results are going to be almost identical.
6:51Because they are physically tethered together. Like they're part of the same long chunk of DNA. Exactly. The tests are highly correlated along the chromosome. If you apply the standard GWS threshold to this data.
7:04You are treating all those tests as if they were completely independent. Which punishes the data far too harshly. Way too harshly. It pushes the threshold so high that you actually destroy your ability to find real evolutionary signals.
7:17So what do previous studies do? Many just bypass this by making up arbitrary ad hoc rules of thumb. They would say, well, if the signal is 1st standard deviations above the mean, we'll just call it evolution.
7:28But using an arbitrary threshold just lets false positive slip through. I mean, falsely concluding that a genetic region is under strong recent selection isn't just a quirky math error. It completely distorts our fundamental map of human biology.
7:42Which is why we need a way to actually control the family-wise error rate. Right. The family-wise error rate. That's the probability of making even one single false discovery across the entire genome scan, right?
7:54Spot on. So to untangle this highly correlated genomic data, the researchers decided to model the standardized IBD rates mathematically. They define the spread of these shared DNA segments as an Ornstein Ulembic process or an OU process.
8:11This is a stochastic model. Meaning it deals with random variables over physical space along the chromosome. Exactly. Here's where it gets really interesting, is the OU process, sort of like tracking a dog on a long leash.
8:24It can wander randomly to the left or right, but there's always a mathematical pole, keeping it tethered, which perfectly mimics how genetic correlations naturally decay as you move further down the chromosome.
8:35Oh, that is a brilliant way to picture it. Yes. The dog wandering represents the random variation in how DNA is shared between people. And the leash. The leash represents the biological reality of recombination.
8:47As you move further down the chromosome, the genetic shuffling that happens generation after generation breaks up those long IBD segments. So the correlation naturally decays. It decays exponentially, yeah.
9:00So the researchers figured out how to estimate the exact tension on that leash, the exponential decay parameter, and they used it to calculate a valid, rigorous significance threshold. Wow. And they develop 2 innovative methods to set these new rules, right?
9:15Oh, they did. First, they used an analytical approximation, which applies a really complex mathematical formula originally developed by statisticians Siegmund and Yakir. And second, they built a computationally efficient whole genome simulation approach.
9:29The simulation approach is wild to me. It's particularly elegant. They generated simulated evolutionary histories for tens of thousands of chromosomes. So instead of waiting 1000s of years to see how humans mate and pass on DNA, the software just fast forwards and rewinds the mathematical probability of human reproduction.
9:47Exactly. They created terabytes of simulated genomic data, where they controlled all the variables just to prove their new thresholds worked. And the results were highly robust. Incredibly robust. By applying this new mathematical leash, they demonstrated greater than 50% statistical power to detect what geneticists call hard selective sweeps.
10:10A hard sweep being an instance where a single beneficial mutation rapidly rises in frequency. Right. They could detect these as long as the selection coefficient was greater than or equal to 0.01. Now, a selection coefficient of 0.01.
10:25basically translates to individuals with the trade, having a one% reproductive advantage over those without it. Yeah. And while one% sounds negligible in our daily lives, in evolutionary terms, compounding over centuries, a one% advantage acts like a tidal wave, it causes a trait to dominate a population remarkably fast.
10:44So they proved the math worked in a pristine simulation, but the ultimate test is applying it to messy, real world human data. Oh, so what happened when they took their strict new multiple testing corrections and applied them to real genomes?
10:58They use the TopiMed project and the UK biobank, right? Yes. Massive databases? They looked at individuals with African and European ancestry from the US, as well as white British, Indian, British, and black British populations.
11:13And how did the new math change the landscape of the data? Drastically? For genetic segments, longer than 2.0 sent to Morgan. Which is just a unit research is used to measure genetic distance in the likelihood of recombination, right?
11:26Right, exactly. So for those segments, the new analytical threshold was calculated at around 2 times 10 to the -6. This new rule acted like a massive industrial filter. It wiped out dozens of signals that older, looser ad hoc rules had previously flagged as significant evolutionary events.
11:44Wow, it's the equivalent of wiping a smudge off your telescope lens and realizing half the stars you thought you discovered were literally just dust. That's exactly what it was like. But crucially, the true evolutionary signals survived the filter.
11:57Oh, so the real ones made it through. Yeah. In the European and Indian ancestry groups, the model successfully verified the massive signal at the LCT gene. That's the exact gene we discussed earlier, the one that controls lactase persistence.
12:10The milk chain, yeah. The filter also preserved a heavy signal at OCA2, which is a gene fundamentally involved in human pigmentation. And we know from ancient DNA sequencing that those are real, confirmed evolutionary adaptations. Exactly.
12:25And the filter worked just as well in the African ancestry group. What did they find there? They found significant signals at the HBB gene complex, which is responsible for hemoglobin. Mutations in that specific region are famously linked to malaria resistance.
12:39Which is a massive life-saving evolutionary sweep. Absolutely. They also verified signals at SEMA 5A, a gene involved in the integrity of the optic nerve, and test 2R1. TS2R1. That belongs to a family of bitter taste receptors, right?
12:54The ones that help humans detect toxins in wild plants. You've got it. Finding those known, validated signals proves that the mathematical leash isn't too tight. It lets the real discoveries through. Right.
13:06But what is truly fascinating to me is what happened when they looked at the largest, most undeniable spikes in the data? Oh, this is the best part. They found massive genomic signals on chromosome 16Q 12.3 near a gene called XYLT1, and another massive spike on chromosome, 22Q11.21.
13:24Yeah. And those massive signals weren't isolated to just one group. They appeared across almost all the different ancestry groups analyzed. So by the old loose rules of evolutionary biology, this would be hailed as a monumental discovery of a shared recent human adaptation.
13:40It would look like the ultimate global survival trait, but under the strict lens of this new mathematical framework. And when cross referencing with how the physical genome is actually structured, the entire story falls apart.
13:51These aren't examples of recent human adaptation at all. So what does this all mean for the data? Are you saying a literal structural glitch in the gene on, like, a missing physical chunk of DNA can artificially inflate the shared IBD segments and completely mimic the exact signature of a life-saving evolutionary adaptation?
14:09That is precisely the trap geneticists have been falling into. Unbelievable. The researchers realize that these massive signals correspond directly to highly variable regions of the genome known for structural variants, specifically recurrent deletions.
14:24Right, like chromosome 22 Q11.21. That region is famous for a structural deletion that causes DeGeorge Syndrome. Exactly. In these regions, a large chunk of physical DNA is simply missing in a significant portion of the population.
14:38Okay, so imagine 2 people handing you a copy of the exact same book. But in both copies, chapters 4 through 10 have just been ripped out. If you only look at the scene where chapter 3 meets chapter 11.
14:51The books look perfectly continuous and identical. Yes. And the sequencing algorithms that read our DNA act just like someone reading that torn book. When a massive chunk is missing, the algorithm just stitches the remaining ends together.
15:03So the software interprets that false continuity as a massive identity by descent segment. Exactly. It sees a bioinformatic mirage. To the algorithm, it looks exactly like a selective sweep where a beneficial trait rapidly spread.
15:15When really the lack of genetic variation isn't because of the survival of the fittest, it's just because the sequence literally isn't there to have variation in the 1st place. It's wild. The dark matter of the genome, these physical, structural glitches are literally impersonating our evolutionary history.
15:33Which really brings us to the broader implications of this deep dive. Why does the broader scientific community need to care about this distinction right now? Because this completely rewrites the landscape of genetic discovery.
15:46The research definitively proves that failing to properly adjust for multiple testing when scanning for shared DNA segments is a massive vulnerability in evolutionary biology. So if researchers continue to borrow outdated thresholds from GW's or just rely on heuristic rules of thumb.
16:02They are going to publish false claims. We will see papers theorizing about the incredible adaptive benefits of a sequence of DNA that is actually just a sequencing black spot. Right. So moving forward, the field really has to adopt these new analytical and simulation-based thresholds.
16:18They do. And it also means researchers need to take this new tool and fundamentally reevaluate structural variance globally. There is an urgent need to go back through the existing literature and ask how many previously reported evolutionary sweeps were actually just structural deletions.
16:35But the researchers are transparent about the limitations of this new mathematical model, too. They are. The Ornstein Ullenbeck process, as applied here, assumes that you are working with large sample sizes at least a few thousand individuals.
16:47Okay. It also assumes you were looking at large panmic dick populations. Pan mictic simply meaning a population that mates completely randomly. But wait, humans rarely mate completely randomly in massive global pools.
17:01We have incredibly complex histories of migration and isolation. Doesn't that mean this math still has a major blind spot for small founder populations? You've identified the exact boundary where this current math requires further refinement.
17:13Yeah. When you have a small founder population or a population that has survived a severe historical bottleneck, the distribution of shared DNA gets incredibly clumpy and chaotic due to genetic drift. Right, the math gets messy.
17:27Exactly. In those isolated groups, the smooth, predictable bell curve assumed by the OU model starts to break down. For those specific scenarios, geneticists will need to develop even more complex mathematical frameworks to handle the extreme variants.
17:41But for the vast majority of large scale global populations, this current framework is a massive leap forward. Oh, absolutely. By applying a rigorous mathematical model to how identical DNA segments are distributed across chromosomes, researchers can finally separate true recent genetic adaptations from structural mirages.
18:01This crucial statistical reality check ensures that our map of human evolution is built on solid, verifiable data rather than mathematical noise. It really forces the entire field to take a step back and question the patterns we thought we understood.
18:15It's a necessary scientific reality check. What does this mean for the countless published studies of genetic adaptation that never accounted for this kind of statistical noise, and how many textbook examples of human evolution might actually just be structural glitches?
18:31Yeah, that is something the whole field has to grapple with now. Because sometimes when we look up and think we're seeing a grand cosmic constellation of human survival. We really are just letting our brains connect random missing dots in the dark.
18:43Exactly. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating.
18:58If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track, created especially for this episode and inspired by the article you've just heard about.
19:07Thanks for listening, and join us next time as we explore more science. Base by base. Late night screens in a silent skin lines across the genome, like a moving map. Every spite looks like a promised land, but shadows rhyme with signal. In the gaps, so we measure the drift between the beats, how nearby test leaning, never alone.
19:58If the echoes talk, the math must speak. Set the bar where truth can hold its own. Draw the threshold in the noise, let it glow. Too low, chase a 1000 ghosts. Keep the family wise fire under control till the real rare signal Matters most and when the peaks all shimmer side by side, we ask is it selection.
20:35Or a hidden slide. IBD threads, long segments we can trust standardize rates like weather on a wire. A process with a memory in the dust got a correlated bending, but not tired. Someone does run hot, some cool too soon.
21:08Choice segments, cheat, conservative truth, bloom, power. Power rises when the sweep is strong and clear, but correction turns the store is near. Shed bright regions. Strange familiar clues. Might be structural.
21:38The skin can't lose. Draw the threshold and the noise that glow. Not too low to chase a thousand ghosts. Family wise fire under control till the real rare signal matters most under drifting skies where correlations hide, we lift the light.
22:10So discovery survives.