A statistical framework uses large reference allele-frequency data (ExAC) together with disease prevalence, heterogeneity, penetrance, and sampling variance to set rigorous frequency filters that improve Mendelian variant interpretation.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and read us in your podcast app. You know, when we talk about human genetics, it is so easy to picture our DNA as this pristine, flawless blueprint.
0:19Oh, absolutely. The textbook image, right? Right. But the reality is actually far messier. I mean, every single one of you listening right now is walking around with roughly 12,000 to 14,000 genetic variants that actively alter your proteins.
0:32Your genome is just packed with mutations. It is entirely true. And, you know, the vast majority of those variants are completely harmless bystanders. Yeah. They make you who you are. They account for human diversity, but they do not cause disease.
0:46Right, but imagine for a second that you're a patient with an undiagnosed, severe genetic condition. You go to the doctor and they sequence your DNA to figure out what's wrong. Which is a standard step nowadays.
0:56Exactly. But how do they find the one specific mutation causing your disease hidden inside that massive microscopic haystack of 14,000 harmless variants? A huge problem. It is. And historically, clinical genetics relied on, well, arbitrary rules of foam to filter that list down.
1:13The general assumption in the medical field was basically this. If a variant is present in less than one% or maybe one in a 1000 people in the general population. It's rare enough that it might be the culprit.
1:26Which, uh, if you're just glancing at the problem, sounds somewhat reasonable, you assume rare diseases are caused by rare mutations. Right, but what if that filter is completely flawed? That sparks a massive question for our deep dive today.
1:40What really happens when the frequency cutoffs we rely on to diagnose severe diseases are actually too lenient. Yeah, that's the big one Could we be actively misdiagnosing totally harmless variants as the underlying cause of rare life altering diseases?
1:54The short answer is yes. And fixing that profound problem requires a complete shift in how we interpret the human genome. So, today we celebrate the work of this international research team who have advanced our understanding of clinical genome interpretation.
2:08We're looking at a landmark paper led by researchers Nicola Wiffin, Eric Medical, James S. Ware, and an extensive collaborative team. Yeah, it was a huge effort. across institutions like Imperial College London, the Broad Institute of MIT and Harvard, and Massachusetts General Hospital.
2:26Incredible team. Truly. And the study is titled Using High Resolution Variant Frequencies to Empower Clinical Genome Interpretation. It was published as an advanced online publication on May 18, 2017. And that was in genetics in medicine, right?
2:40Yes, exactly. The official Journal of the American College of Medical Genetics and Genomics. Awesome. So to really grasp why this paper was such a monumental game changer for medicine. We have to look at the clinical dilemma geneticists were facing at the time.
2:53Right. The technology was moving faster than the interpretation. Exactly. Technologies like whole exome and whole genome sequencing had absolutely revolutionized our ability to discover the genes tied to mendelian diseases.
3:05Those being the conditions caused by mutation in a single gene. Right, like cystic fibrosis or Huntington's disease. But distinguishing the pathogenic, the disease causing variants from the benign bystanders, that remained a daunting, almost paralyzing challenge for doctors.
3:23It's the ultimate signal to noise problem. I mean, you have incredible sequencing technology generating mountains of data for a single patient. But having a spreadsheet full of data is not the same as having a diagnosis.
3:36No, definitely not. And the clinical standard had been using those lenient cutoffs you mentioned. Laboratories were looking at a patient's DNA and saying, well, this variant is found in less than one% of people, so it's a candidate for recessive disease.
3:49Or it's found in less than one in a 1000 people. So it's a candidate for a dominant disease. Exactly. That was the rule book. Okay, let's unpack this because those old rules of thumb, just do not make intuitive sense when you really think about them.
4:02It's kind of like, um, setting a height requirement for a roller coaster, but setting it so low that toddlers can get on. That's great way to put it. Right. Because if a severe disease only affects, say, one in 10,000 people globally, how could the mutation causing it be present in one out of every 100 people?
4:21It just mathematically doesn't work. Exactly. The math of that old standard completely falls apart. You're letting 1000s of harmless variants slip through the filter and land right in the suspect pile.
4:32What's fascinating here is that the field of population genetics has always fundamentally disagreed with those lenient cutoffs. Really? So the scientists knew it was off. Oh, yeah. The math of human evolution simply dictates that severe disease causing variants must be much, much rarer than one percent.
4:52I mean, think about the mechanics of natural selection. Right, purifying selection. Exactly. If a mutation causes a severe disease that heavily impacts your health, people carrying that mutation are far less likely to pass it on to the next generation.
5:05So the mutation essentially just weeds itself out of the population over time. Precisely. The only real exceptions are when you have very specific circumstances, like, a bottlenecked population where a small group of people expanded rapidly, and a specific mutation got trapped in that lineage.
5:21Oh, like founder effects? Yeah, exactly. Or situations evolving balancing selection. That's where carrying one copy of a bad mutation actually gives you an evolutionary advantage. Like how the sickle cell trait provides some resistance to malaria.
5:35Spot on. But outside of those extremely rare exceptions, severe disease variants simply do not hang around in the general human population at high frequencies. So if the scientific consensus knew these cutoffs were wrong, why were clinics still using them?
5:50Because for a long time, the field simply didn't have the data to do anything better. I mean, they urgently needed a more stringent, statistically robust approach. They needed a better filter. Right. We needed a framework that actually accounted for the specific genetic architecture of individual diseases rather than just slapping a one-size fits all percentage across every single condition in the medical textbook.
6:11Which makes sense. But to build a stricter filter, you need to know exactly what normal looks like across the entire human race. Which brings us to the monumental effort that made this whole shift possible.
6:23Because to build this framework, the researchers needed an unprecedented amount of genetic data. They really did. They needed the exome aggregation consortium, better known as exile. Zazi was revolutionary.
6:37It was this massive database detailing the allegal frequencies of 10000000 genomic variant. 10 million. 10 million. Pooled from over 60,000 humans. And before exazi, our reference data sets were much smaller, maybe a few 1000 people at best.
6:52Wow, that's a huge jump. Huge. And when your baseline is that small. You cannot accurately estimate the frequencies of incredibly rare variants. You just don't have the sample size to say anything meaningful.
7:03Well, you'd never even see the really rare stuff. Exactly. But by aggregating the exomes of 60,000 individuals, we finally had the statistical power to get robust, reliable frequency estimates. So, armed with that massive database.
7:17The research team created a two-stage statistical framework. Instead of saying, let's just filter out anything more common than one in a 1000, They decided to calculate a maximum credible population alleal frequency.
7:31Which was specifically tailored to whatever disease the doctor was looking at. Exactly. So what does this all mean? Think about it like household budgeting. You know your absolute maximum budget for rent based on your specific monthly income.
7:45If you make $4000 a month, your rent cannot mathematically be $5000. Unless you're going into serious debt. Fair point, but you get the idea. Similarly, a pathogenic variant simply cannot be more frequent in the general population than the actual prevalence of the disease it's responsible for.
8:03It's an absolute upper limit. That budgeting analogy captures it perfectly. You cannot spend more than you have, and a variant cannot be more common than the disease it causes. And to set that accurate budget.
8:14The team incorporated 3 brilliant variables to refine their mathematical rule. The first, as you mentioned, is disease prevalence. Literally, how common is the condition in the general population? Is it one in 500 or is it one in a 1000000?
8:30Exactly. But they didn't stop there. Because a disease isn't always a one-to-one relationship with a single mutation. They also factored in a concept called heterogeneity. Right. genetic and allelic heterogeneity.
8:43Yeah. This asks the question. What proportion of cases of this specific disease are actually caused by the single variant we're looking at? Because there could be dozens of causes. Right. If a disease is caused by 50 different mutations scattered across 10 different genes, no single mutation is going to be responsible for 100% of the patients.
9:04So if the variant you're studying only accounts for 5% of all cases of the disease, then its budget in the general population has to be restricted to reflect just that 5% slice of the pie. That makes perfect sense.
9:15It's a proportional limit. And the 3rd crucial variable they wove into this equation was penetrance. Oh, penetrance. That's a big one It is perhaps the trickiest but most important piece of the puzzle.
9:27Penetrance asks. If you carry this specific variant in your DNA, what are the actual odds, you will develop the disease? Because having a mutation doesn't always guarantee you get sick. Exactly. Some mutations are fully penetrant.
9:43If you have the mutation, you will absolutely get the disease. But other mutations have reduced penetrance, meaning you might carry the mutation, but you only have, say, a 50% chance of ever showing symptoms.
9:54But why does that change the math so much? Because of camouflage. If a variant has a low penetrance, it can hide in perfectly healthy people. If only 10% of the people who carry a variant actually gets sick.
10:07That means 90% of the carriers are walking around healthy. and passing the variant onto their children. Exactly. Because it's hiding in healthy carriers, the variant can be much more common in the general population without triggering a massive spike in the overall prevalence of the disease.
10:22Okay, I see. So they take prevalence, heterogeneity and tenetrants, and they weave them together to find the absolute maximum credible frequency for a specific disease variant. Yes. But hold on. I have to jump in here with a reality check on the data.
10:3660,000 people in the exact sea database is a lot of people. It was a herculean effort. Absolutely. But there are 1000000000s of people on Earth. 60,000 is still just a tiny fraction of humanity. How do they know that what they see in excessi isn't just a fluke.
10:53That is the exact hurdle the team had to cross next. How do you account for the randomness of sampling? Right, because random chance is a huge factor. It is. If a variant is incredibly rare, seeing it 3 times in 60,000 people versus seeing it 4 times, could just be the luck of the draw.
11:10It doesn't necessarily mean the variant is more common globally. It might just mean you happen to sample 3 people from the same town. Exactly. To handle that sampling variants. The team took a highly innovative mathematical step.
11:21They used a poisson probability distribution to calculate what they called the maximum tolerated allegal count. Okay, a Poisson distribution. Yeah, it's basically setting a hard, absolute limit on how many times a disease causing variant can show up in the exact database before you can confidently say, nope, that's too common.
11:40It cannot be the cause. So it's about statistics. Yes. And instead of using a textbook definition, Let's think about how the Poisson distribution works in the real world. Imagine, you know, a certain traffic intersection averages exactly one car crash a year based on decades of data.
11:56Okay. The Poisson distribution is the math that tells you the exact odds of seeing 3 crashes there next year just by pure, terrible luck. It models the probability of random events. I like that analogy.
12:08Thanks. The researchers use the same math to ask. What are the odds that a variant is truly rare enough globally to cause this disease, but just happen to show up a dozen times in our group of 60,000 by pure chance?
12:20Ah, so they used it to calculate a confidence interval. They could say, with 95% certainty, a true disease causing variant for this specific condition should not appear more than X number of times in the exact C database.
12:33Exactly. They built the theoretical framework and then they road tested it. And their primary real-world test case was hypertrophic cardiomyopathy or HCM. Which is a brilliant choice. It really is. HCM is a dominant cardiac disorder that causes the heart muscle to become abnormally thick.
12:49It can lead to severe complications, including sudden cardiac arrest. And it affects about one in every 500 people. Right. So it's relatively common for a genetic disorder. Yes. And it's genetic architecture, it's incredibly well studied.
13:03We have a lot of data on it. So let's walk through the exact math they used because it beautifully illustrates the power of this new filter. Let's do it They knew the prevalence of ACM was one in 500. Then they looked at the most common known genetic variant that causes the disease, which sits on the MYBPC3 gene.
13:21Okay. Based on decades of clinical data, that single variant is responsible for about 2% of all HCM cases. So the heterogeneity factor, its slice of the pie is 2%. Right. Next, they looked at penetrance.
13:35Based on clinical reports for that specific variant, they assumed a penetrance of 50%. Got it. When they plugged all of those numbers, the prevalence, the 2% contribution, the 50% penetrance, into their new framework, the math calculated a maximum credible population frequency of about 4 in 100,000.
13:53About 4 and 100,000. That is incredibly rare. much, much rarer than the old one in a 1000 rule of thumb. It is. But here is the critical step. They map that frequency onto the exact C database. At the time, exact seat contained about 121,000 sequenced chromosomes for this specific gene.
14:12Using the puss on distribution. The framework translated that frequency of 4 in 100,000 into a maximum tolerated alleal count. Out of 121,000 chances, What is the absolute maximum number of times this variant can appear.
14:27And what was it? The answer was exactly nine. Here's where it gets really interesting. Just let that sink in. Nine. Out of 121,000 chromosomes, if a variant shows up 10 times, it's mathematically too common to be the sole cause of hypertrophic cardiomyopathy.
14:44That is a stunningly strict threshold, especially when you realize the old standard would have allowed that same variant to appear over a 100 times before anyone batted an eye. This trictness is the point.
14:56But, you know, it also presents a risk. If you make your colander's holes too small, you might accidentally cash the sand along with the gold. It's a good point. The critical test was, did this strict limit of 9 accidentally throw out the true disease causing mutations?
15:12To find out, the team applied this threshold to a massive data set of over 6000 published verified HCM cases, and the results were phenomenal, 99.6% of the known truly pathogenic variants, appropriately fell below that threshold of nine.
15:28Which proves that you could be mathematically rigorous without losing the true pathogenic variants. The clinical diagnostic power of this finding cannot be overstated. I mean, imagine a doctor sitting down with an average patient's exum sequencing results.
15:43They're looking at a massive list of potential candidate variants, trying to find the one causing the patient's heart to fail. It's overwhelming. By applying this new strict threshold, the framework reduces the number of candidate variants a doctor has to manually review by two thirds.
16:00It drops the list from an average of 176 variants down to just 63. It literally shrinks the haystack by two thirds, all without falsely throwing out the true disease causing mutations. The false positive rate was less than one in a 1000 But the impact didn't stop at just making diagnoses faster.
16:18It also allowed the team to go back and correct the historical record. Right. Looking at past data. They look at Clinvar. Clinvar is a massive public archive where genetic testing laboratories report relationships between specific variants and diseases.
16:31Using this new math. The framework caught over 40 variants that had been previously labeled by laboratories as pathogenic or likely pathogenic for HCM, but were actually appearing way too often in the exassi database to possibly be causing the disease.
16:47Wow. Stop and think about the human impact of that. Over 40 variants that doctors might have been looking at thinking, ah, here is the fatal cause of my patient's heart condition. Exactly. They might have been advising patients against playing sports or screening family members, causing immense anxiety.
17:04When in reality, the math proves those variants were just innocent bystanders. It's heartbreaking to think about, but also so crucial to fix. Because of this study, those variants were recurated and reclassified as benign, or as variants of unknown significance.
17:19That is a massive course correction that directly impacts patient care. If we connect this to the bigger picture, the beauty of this statistical framework is that it's not just for hypertrophic cardiomyopathy, it is universally adaptable.
17:32Oh, to any genetic disease. Pretty much. The researchers demonstrated its application to other dominant conditions, like Morphon syndrome, a connective tissue disorder that can cause severe cardiovascular problems, and Noonan syndrome, which affects development, and crucially, the framework easily adapts to recessive diseases as well.
17:51Recessive diseases are really interesting because the math fundamentally changes. With a recessive disease, you need 2 copies of the variant to actually get sick. So if you only have one copy, you're a healthy carrier.
18:04Because of that, these variants can hide much more easily and float around at significantly higher frequencies in the general population. Exactly. For recessive disease, you actually have to use the square root of the disease prevalence to find the maximum allegal frequency.
18:18That sounds complicated. It does, but the framework handles it beautifully. The team tested this on primary soluri dyskinesia or PCD, which is a severe recessive disorder affecting the lungs and respiratory tract.
18:30Okay. When they ran the specific genetic architecture of PCD through the framework, It's set a maximum tolerated exact C count of 322. 322. And when they applied that limit to the Clinvar database, the framework immediately flagged one supposedly pathogenic variant for PCD that appeared over 2300 times an exact scene.
18:53Insane, right? It was an absolute smoking gun of innocence. The math proved beyond a shadow of a doubt that this variant was way too common to be the sole cause of the disease. And what is truly remarkable about this research team is that they went far beyond just publishing a brilliant theoretical paper.
19:10They wanted to fundamentally change clinical practice on the ground. So they didn't just leave it in the journal. No, they precomputed these filtering values for all variants across the entire exact z database.
19:21They literally did the math for everyone. Yep, Yep. they built an online calculator. So a clinician sitting in a hospital anywhere in the world could just plug in the prevalence, the heterogeneity and the penetrants of whatever disease they're studying.
19:35Then the tool immediately tells them the maximum credible frequency. It directly empowers clinicians worldwide to filter out the noise and focus on the truth threats. But I do have to push back a little here, because nothing in biology is ever perfectly neat.
19:48Fair enough. This framework is incredibly elegant when you know the variables. When you're looking at well studied diseases like HCM, it works like a charm. But what if a disease isn't well characterized?
20:00That's a valid point. What if we genuinely do not know the exact penetrance of a rare condition? Or we have absolutely no idea how many different genes are involved in its heterogeneity. Doesn't the math fall apart if the inputs are just educated guesses?
20:13That is a very fair critique. And it points to the primary limitations of this study. The framework does rely heavily on our current understanding of a disease's architecture. And as we discussed earlier, penetrance is notoriously difficult to estimate accurately.
20:29If you assume a variant has a 50% penetrance. But the true penetrance of this newly discovered variant is actually only 5%. Your maximum frequency threshold is gonna be way too strict. Right, because a very low penetrance means the variant is incredibly good at hiding in the population without causing disease, so it should be allowed a higher budget or frequency limit.
20:51Exactly. The researchers openly acknowledge this limitation. However, they argue that variants with extremely low penetrance often have questionable diagnostic utility in a clinical setting anyway. That makes sense. Because if a variant only causes disease in one out of every 100 people who carry it, finding it in a patient's DNA doesn't really give the doctor a definitive answer.
21:13You're still left guessing. Exactly. But to address the uncertainty of these variables. Their online calculator actually allows users to define a broad range of compatible penetrances rather than a single fixed number.
21:25Clinicians can explore different architectural scenarios and see how the limits change. Okay. That helps. But what about the exact T database itself? The whole framework rests on exact scene. We're assuming it represents a perfectly healthy baseline population to compare our sick patients against, but is that entirely true?
21:42This raises an important question about the nature of our reference data sets. Exact days is not a perfectly healthy cohort. It was aggregated from dozens of various research studies, some of which specifically included patients with complex diseases, like type 2 diabetes, schizophrenia, or coronary artery disease.
22:01Oh okay. So while it's generally depleted of severe childhood mendelian conditions, because those patients usually were included in these adult studies, it might be slightly enriched for common adult diseases.
22:14Meaning, if you're using this framework to study a genetic mutation linked to something like early onset diabetes, you have to be very careful because your normal baseline might actually be packed with people who have that exact treat.
22:26Precisely. You cannot just blindly trust the algorithm. This mathematical tool must be used alongside other clinical data. You still need to look at the whole picture. Exactly. You still need to look at amino acid conservation.
22:38Is this variant disrupting a crucial part of the protein that hasn't changed in 1000000s of years of evolution? You still need to look at functional data in the lab. You still need to look at family segregation.
22:49Does the variant actually track with the disease within a specific family tree? So it's not a magic bullet? No. The allele frequency framework provides a powerful necessary constraint, but it's still just one piece of a much larger diagnostic puzzle.
23:07So summarizing this entire deep dive into just a few sentences. By abandoning arbitrary percentage cutoffs and adopting a mathematically rigorous disease specific statistical framework, clinical geneticists can effectively filter out the harmless variants clogging up our data.
23:23Which is huge. It dramatically reduces diagnostic noise, saves patients from agonizing misdiagnoses, and does it all without losing the true pathogenic mutations hidden in our genomes. It replaces a blunt instrument with a surgical tool.
23:36And it sets an entirely new standard for how we handle the absolute flood of genomic data that modern medicine is producing. Which leaves us with a massive thought to chew on. We've spent this entire conversation talking about finding the one mutation in a haystack of 14,000 variants using a database of 60,000 people.
23:54But our data sets are expanding exponentially. What does this mean for the future of personalized medicine at our reference databases grow from 10s of 1000s to 1000000s of individuals. That's the question we have to answer next.
24:06This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating.
24:21If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
24:30Thanks for listening, and join us next time as we explore more science base by base.