This episode examines a theoretical and empirical study showing how the geographic breadth of sampling affects discovery and observed frequencies of deleterious rare variants. The authors develop a spatial stochastic model, validate it with simulations, and test predictions using UK Biobank exome resampling.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine you're hunting for a genetic mutation that causes a rare heart disease.
0:13If you test 100,000 people in a single city, you might actually find it. Right, because of the local concentration. Exactly. But if you test 100,000 people spread uniformly across an entire country, you might completely miss it.
0:27Which sounds completely backwards, I know. It really does. But why does that happen? Because when it comes to human DNA, you know, casting a wider net actually changes the reality of what you catch. We are living through an absolute explosion in the scale of human genetic data right now.
0:44Oh, absolutely. I mean, human biobanks, these massive repositories of genomic information, they're just growing exponentially. We're looking at data sets now, reaching 100s of 1000s of individuals. Right. Right.
0:55But as we build these massive databases, we're also dramatically expanding where we sample these individuals from. We're shifting from like localized city scale samples to national or even global scales.
1:06Yeah, which brings up the core issue. Right. So what really happens when we cast a wider geographic net for human DNA. Does it actually change the genetic variants we discover? It absolutely alters the genetic landscape we uncover.
1:20When you change the geographic breadth of a study, you know, the physical area you're drawing your subjects from, you fundamentally alter both the frequency and the sheer number of the genetic variants you discovered.
1:30Wow. Yeah, it's a profound shift in perspective that essentially dictates what we can and cannot see. Which means this simple question of geography could completely change our hunt for disease causing genetic mutations.
1:43It really could. And today, we celebrate the work of Margaret C. Steiner, Daniel P. Rice, John November, and their colleagues across the University of Chicago, MIT, Secure Bio, Johns Hopkins, and the Icon School of Medicine, who have advanced our understanding of study design and the sampling of deleterious rare variants in buyer bank scale data sets.
2:04That is an absolute powerhouse of a research team. It really is. And their work was published in the journal PNAS in June of 2025. Okay, let's unpack this. Before we dive into the dense math and the biobanks of today.
2:15We should probably, you know, set the stage for why this specific problem matters so much to researchers right now. Yeah, setting the historical context is crucial here. And where this whole line of questioning actually started.
2:30Because in human genetics, researchers are practically desperate to find rare large effect variants, right? Exactly. We're talking about the specific mutations that actually cause complex diseases or significantly increase your susceptibility to them.
2:45Finding those mutations is basically the holy grail of precision medicine. But from an evolutionary standpoint, the universe is actively working against us finding them. Because of natural selection. Yes, natural selection, fiercely weeds out those kinds of harmful mutations through a process called negative selection.
3:04Right, okay. Because these mutations are deleterious, meaning they are harmful to the organism's ability to survive and reproduce evolution, continuously pushes them down to very, very low frequencies in the population.
3:15So they don't disappear entirely then. No, they don't disappear entirely simply because new mutations are always randomly springing up, but they do remain incredibly rare. I'm trying to connect the dots here, though, because looking at the sources, this spatial puzzle like how rare variants are physically scattered across a landscape, it isn't any question at all.
3:35No, not at all. It goes back decades. Right. We got pioneering geneticists like Sewell Wright and Theodosius Dibzansky in the 1940s, and later Bruce Wallace in the 1960s, studying this exact spatial phenomenon.
3:48But they weren't looking at human biobanks. No, they were not. I mean, they were chasing fruit flies. Yeah, the classic Drosophilus studies. Exactly. Wallace actually figured out that the rate at which you'd find distinct lethal mutations in Drosophila decayed exponentially based on the geographic distance between the fly populations he sampled.
4:08They did, yeah. So how does hunting for lethal traits in 1940s food flies, conceptually map onto modern biobanks, looking for complex human disease traits? Just feels like 2 completely different worlds.
4:20I get that, but the jump from fruit flies to human biobanks makes perfect sense when you realize that the underlying spatial mathematics of inheritance, mutation, and geography. They remain exactly the same.
4:32Wait, really, the math is identical. Essentially, yes. Those early researchers were trying to understand how a lethal mutation physically moves through a natural environment over time. They were, you know, literally walking between different fruit fly habitats measuring distance.
4:48Today we face the exact same spatial sampling questions. The only difference is that instead of walking through fields with a net, we're querying massive computational databases of human DNA. Oh, that makes a lot of sense.
5:01Yeah, we are looking for the human equivalence of those fruit fly mutations, rare variants that cause disease. And we desperately need to know if the geographic way we collected our human samples is distorting our view of how common those diseases actually are.
5:15Okay, so we know from the fruit flies that geography hides these mutations, but you can't just walk across the UK with a butterfly net collecting human DNA. No, definitely not. How on earth did they model that mathematically to bridge the gap to modern humans?
5:29Well, the research team used a highly rigorous, three pronged methodological approach to solve this. The 1st prong was building a purely theoretical mathematical model. Okay, lay it on me. They formulated what is called a stochastic branching process in continuous space.
5:46Uh, a stochastic branching process sounds like a tool a Wall Street hedge fund uses to predict the stock market. Well, actually, it's very similar in its use of probability. Think of it as a way to calculate the odds of a family tree spreading across a map over 1000s of years.
6:02Okay, I can picture that. It's a mathematical environment that simultaneously accounts for 3 things. First, a new mutation randomly appearing, that's the starcastic part. Second, the negative selection, constantly pruning that family tree.
6:15And third, how that mutation physically diffuses or spreads over a two-dimensional geographic area generation by generation. And the crucial parameter they introduced in this math model is something they call W, right?
6:28Which measures the sampling breath. Exactly. Parameter W. To help you visualize this. Imagine you're looking at a map of a country in a dark room. The sampling breath. WI is like a spotlight you drop onto that map.
6:42You love that analogy. A very narrow sample. So a small levy is like a concentrated 50 kilometer laser beam focused on a single city. You are only seeing the genetic variants illuminated within that tiny intense beam.
6:55Right. But a broad sample is like swapping the laser beam for a uniform diffuse floodlight that washes over the entire country. That is a brilliant way to picture it. And to prove that their math regarding the laser beam and the floodlight actually worked in reality, they moved to the 2nd prong of their methodology.
7:13Which was? Computer simulations, they used a sophisticated software program called Slim to run individual-based forward-time genetic simulations. Wait, let's remind the lister, because this is a crucial detail in the paper.
7:25They didn't just simulate abstract strings of code. They simulated deploy genomes. Yes, very important distinction. When you say diploid. We're talking about the fact that we inherit two sets of chromosomes, one from each parent, right?
7:36How does that complicate the simulation compared to simpler models? It makes it exponentially more complex and much closer to human reality. Simple models often assume every variant is just a rare, independent data point.
7:50Which isn't how biology works. Not at all. By simulating deployed genomes. The researchers force the computer to track how 2 sets of chromosomes interact, recombine, and get passed down through multiple life stages of an organism, all while moving across a geographic grid.
8:06Wow, that's intense. Yeah, it was a rigorous stress test of their theoretical math. Basically proving it holds up even when you introduce the messy realities of sexual reproduction and chromosomal inheritance.
8:18But they didn't just leave it in the computer. They took it to the real world for the 3rd prong. They did. They turned to empirical real-world data from the UK Biobank. Specifically, they performed in silica resampling.
8:29And that basically means they computationally created artificial study cohorts using real human ex zone data, right? Exactly. They focus heavily on the protein coating regions of the genome, specifically chromosome one.
8:42And because the UK Biobank actually includes the birthplace coordinates for its participants, The researchers could literally simulate our spotlight metaphor. They could. It's perfectly set up for that.
8:54I love this part. They could look at a 50 kilometer radius, which, by the way, getting a 50 kilometer genetic sample in the middle of dense, diverse London is going to look a lot different than a 50 kilometer radius in the rural Scottish highlands.
9:07Undoubtedly. But they compare that narrow 50 kilometer sample against a 100 kilometer sample, 150 kilometer sample, and finally, uniform sampling across the entirety of Britain. And that transition scaling up from 50 kilometers to the whole country is exactly what revealed the core phenomenon of the paper.
9:23Here's where it gets really interesting. Like what actually happens when you switch between the narrow laser beam and the wide floodlight. The researchers unveiled a massive fundamental trade-off that they call discovery versus dilution.
9:38Yes, the discovery versus dilution tradeoff. This dictates everything about what a genetic study will find. Let's start with the floodlight, the uniform sample. Okay. If you use broad uniform sampling across the whole country, You trigger the discovery effect.
9:55You end up discovering a much greater total number of distinct genetic variants. Which on the surface sounds fantastic. I mean, more variants mean more potential discoveries for medicine, right? It does sound fantastic until you encounter the dilution effect.
10:09Because you're casting such a wide net, you are capturing 1000s of different localized family histories across a massive area. Oh, I see. So those distinct variants you just discovered are heavily diluted.
10:21Most of the new variants you find will appear as singletons. Singletons. Yeah, meaning out of your massive sample of 100s of 1000s of people, you only see that specific mutation in exactly one person. Oh wow.
10:33So when we switch to the floodlight, we're eliminating way more of the map, but the light hitting any specific mutation is so dim, so diluted that it barely registers as a singleton. What if we swap back to the narrow laser beam?
10:47When you use the narrow laser beam, you find fewer unique variants overall, the total count of distinct mutations drop significantly. However, the variants you do find are highly concentrated. Because of local ancestry.
11:01Right. Because you're sampling people who live physically close together and therefore likely share more recent geographic ancestors, the variants you discover are present at much higher frequencies in your sample.
11:12You see them in dozens or 100s of people, not just as isolated singletons. Let me pause you, because this paper tripped me up hard here. This leads to a crucial mathematical paradox. Yes, it does The researchers state that despite these massive shifts in what you discover, you know, lots of singletons with the floodlight, higher frequencies with the laser beam, the expected average allele frequency across the whole study remains exactly the same.
11:37It is deeply counterintuitive. Yeah. The total expected heterozygosity doesn't change regardless of your sampling breath. Let's define heterozygosity real quick for the listener. There's basically a measure of genetic variation, the probability that 2 alleges chosen at random from the population are different.
11:54That's a perfect definition. So how is it logically possible to find completely different counts of mutations depending on your spotlight? Yet have the overall average genetic variation level remain totally unchanged.
12:05The best way to understand the paradox is to look at the hard data from their UK biobank experiment. When they moved from the narrow 50 kilometer sample to the broad uniform sampling across Britain, they found 72.3% more loss of function variants.
12:21Wait, loss of function variants. Yeah, these are the highly deleterious mutations, the broken genes that often cause severe disease. Finding 72% more broken genes is a staggering increase in raw discovery.
12:33It is massive, but remember the dilution effect. While they found 72% more distinct broken genes, the expected heterozygosity, meaning the frequency at which those specific variant sites appeared in the sample, dropped by 36.75%.
12:47Let me try to summarize this to make sure I'm wrapping my head around the math. If I'm a researcher and I switch to the floodlight, I am basically adding 1000s of new columns to my massive spreadsheet, one for each new broken gene I discovered.
13:02Yes, your spreadsheet gets much, much wider. But because almost all of those new jeans are singletons. I'm just filling those new columns with 1000s of zeros and a single one. Exactly. So while the absolute number of distinct mutations grows wildly, The frequencies of each individual mutation plummet perfectly in tandem, the discovery of new, incredibly rare variants mathematically perfectly cancels out the dilution of the variant frequencies.
13:28You nailed it. So if I average the frequency across every single column on my spreadsheet, that average stays identically constant. Yes. The math fundamentally balances out perfectly. The influx of ultra rare singletons drags the individual site frequencies down at the exact same rate that the total number of sites goes up.
13:45That is wild. But, um, if the overall average stays the same in the spreadsheet, does this spatial scaling actually change how we do science in the real world, like does this change how we actually go about curing diseases?
13:59It changes study design completely, and it impacts 2 major scientific fields in profound ways. Let's look at clinical research first, specifically genome white association studies, or GWYs. Right, GWIs.
14:13These are the foundation of modern genetic medicine. These are the studies where we scan the genomes of 1000s of people define the specific mutations statistically linked to a disease. Exactly. Based on the discovery versus dilution trade-off we just unpacked.
14:26If you build a broad geographically uniform biobank, you will discover far more potential disease targets. You hand researchers a huge list of distinct broken genes. But because of the dilution effect, Those targets appear at much lower frequencies.
14:40Most of them are singletons. And that is a massive hurdle. In a single variant G to wise test, your statistical power to prove a connection is deeply tied to how frequently an allele appears in your sample.
14:50Right? You need a large sample size of the actual mutation. Yes. If a mutation only shows up in one or 2 people out of 100,000, it's statistically incredibly difficult, if not impossible, to definitively prove that that specific mutation cause their disease.
15:06You simply lack the statistical power. So the broad sample gives you a massive map of targets, but strips away your power to investigate them individually. Your trading concentration for discovery. That is the clinical dilemma.
15:19Now, if we connect this to the bigger picture, the 2nd major impact is on evolutionary genetics. Okay, how so? Evolutionary genetics are trying to calculate the true strength of negative natural selection.
15:30They want to know how fiercely evolution is weeding out bad mutations. And how does sampling breath mess up our understanding of evolution? It happens when researchers make a faulty assumption about their data.
15:40Imagine a researcher takes a narrow 50 kilometer sample, just the laser beam on one city. But when they sit down to do the math, they analyze that data as if it represents a random panamictic population.
15:53Let's define canamictic for the listener. That means assuming the population is perfectly and randomly mixed, right? Like a giant blender where anyone is equally likely to mate with anyone else in the country.
16:05Yes, precisely. They treat their single city as a perfect microcosm of the entire global pool, but we know from the narrow laser beam effect that geographically narrow samples artificially concentrate rare variants.
16:20Yeah, because of shared local ancestry. Exactly. They appear at much higher frequencies in that single city than they do globally. If you feed those artificially high concentrated frequencies into your evolutionary models, your math will tell you that these deleterious mutations are much more common worldwide than they actually are.
16:38Ah, I see it now. And if you think harmful mutations are common, you'll fundamentally underestimate how strongly natural selection is fighting to remove them. Right. Because if they were really that common, you would imply natural selections doing a terrible job.
16:51What's fascinating here is that simply ignoring the geographic breadth of your study can lead you to underestimate the strength of negative natural selection entirely. You end up with a skewed understanding of the human motational load, purely because you treated a narrow geographic sample as if it were a well-mixed global pool.
17:10It's incredible how a simple geographic assumption can, you know, ripple out and distort our entire view of human evolution. It really is. But we do have to acknowledge the limitations of the math here, right?
17:22Because the paper points out that this beautiful theoretical model operates on a perfect, infinite, two-dimensional tourist. Yes, the Taurus. Mathematically speaking, a Taurus is essentially a donut shape.
17:34A donut. Okay. What that means for the simulation is that the model has no boundaries. If a simulated human family line migrates off the top edge of the map, they just wrap around to the bottom like Pac-Man.
17:45Right. There are no coastlines or mountains to stop them. But the real world is infinitely messier than a donut. Britain, where they pulled the empirical data is an island. It has hard coastlines. Humans have distinct geographic borders, we have complex histories of rapid population growth, and we have long-range migrations.
18:04Oh, absolutely. I mean, you can hop on a plane in London and be in Tokyo tomorrow. The abstract model certainly does not fully capture those massive leaps in migration or the complexities of modern borders.
18:14In reality, a geographically narrow 50 kilometer sample in a hyperconnected modern city like London might contain genetics from all over the globe, effectively acting like a broad sample in disguise. Right, that makes sense.
18:28But even with those real world complexities, the empirical data from the UK biobank proves that the fundamental mechanics of discovery and dilution still hold true. The math scales, even if the maps are messy.
18:40Well, geographic breath isn't just an administrative detail, then. It is a fundamental knob in genetic study design. It absolutely is. Turning that knob wider reveals more unique genetic variants, but it mercilessly dilutes their frequency.
18:53As a researcher, you're forced to navigate a strict, unavoidable tradeoff between discovery and concentration. It's a delicate balancing act. You cannot cheat the math of spatial distribution, no matter how large your database grows.
19:06So, as we pack up our spotlight and look toward the future, I want to leave you with this to chew on. What does this mean for the design of the next generation of 1000000 person bio banks, and how we balance the need for global genetic diversity, with the statistical power to actually cure disease?
19:21That is the big question. This episode was based on an open access article. Under the CCBY 4.0 license. You can find a direct length of the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app, and leave a 5 star rating.
19:37If you'd like to support our work, use a donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
19:47Thanks for listening and join us next time as we explore more science, base by base.