A biobank-scale study using UK Biobank and Mount Sinai BioMe exomes examines three genetic contributors to incomplete penetrance and variable severity of monogenic cardiometabolic variants: heterogeneous missense variant effects, additive polygenic background, and marginal epistasis between carrier status and common variation.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. You know, we think about genetics.
0:10We uh, we usually picture a very straightforward cause and effect. Right. You inherit a mutated gene, you get the disease. Exactly. It feels completely deterministic. But consider this really surprising fact.
0:21Over 3% of the population carries a dominant disease causing genetic variant what we call a pathogenic variant. Yet only a fraction of those people ever actually develop the disease. Which is pretty mind blowing when you think about it.
0:33It is I mean, why do 2 people with the exact same genetic mutation have completely different health outcomes. You know, one is suffering severe symptoms and the other is just perfectly healthy. What really happens when we look past the single broken gene.
0:47Okay, let's unpack this. Well, today we celebrate the work of Angela Way, Valerie A, our Belita, Noah Zeitlin, and their, uh, their extensive collaborative team across institutions like UCLA, UCSF, Mount Sinai, and Mass General.
1:02That is a massive lineup of talent. Oh, absolutely. And their research, which was published in the journal, Nature Communications in June 2025, has massively advanced our understanding of genetic penetrance and disease severity.
1:15Yeah, and to really understand why this research is such a game changer. I want you to imagine, sitting in a doctor's office. You've just taken a genetic test because you have a family history of high cholesterol or diabetes.
1:29The classic scenario. Right. And the Dr. Lucia chart and says, well, we found a known pathogenic mutation in your DNA. Naturally, your heart just drops. You ask, am I going to get sick? And the doctor just, you know, shrugs.
1:42they can't tell you for sure. And that is the exact clinical problem this deep dive is focusing on. Doctor spot these mutations every single day, but they are constantly running into 2 massive walls. The 1st is called incomplete penetrance.
1:54Which basically means having the mutation doesn't guarantee you'll ever actually get the disease. Precisely. And the 2nd wall is variable expressivity, meaning even if you do get sick, the severity of the symptoms can range from completely unnoticeable to, well, life-threatening.
2:08Right, it's like finding a typo in a recipe. Sometimes the typo says to use a cup of salt instead of a cup of sugar, and it completely ruins the entire cake. Yeah, that would be terrible. But other times, the typo is just in the instructions for how to grease the pan.
2:21You bake the cake, and you don't even notice the error in the final product. I love that analogy. It perfectly captures what's happening. And today, we are focusing specifically on cardio metabolic traits.
2:33Things like high LDL, which is the bad cholesterol, high triglycerides, monogenic obesity, and a specific type of diabetes called MODY. Right. Because for decades, the medical field operated on this concept of monogenic diseases, you know, mono, meaning one, to genic, meaning gene.
2:51The working theory was basically that a single gene mutation held all the power over your destiny for these specific conditions. But the reality, as your recipe analogy suggests, is far more complicated.
3:02Right, it's never just one typo. Exactly. And this study finally provides a really robust framework for understanding exactly why that is. And it does so by analyzing just an unprecedented amount of data.
3:12And when we say massive data, we are talking biobank scale here. The team pulled from the UK biobank. initially looking at over 200,000 ex-somes, and eventually scale that up to over 454,000. That is a staggering number of X-Homes.
3:27It really is. And just to be clear for everyone, when we say X-OMs, we aren't talking about the entire 3000000000 letters of your DNA genome, right? Correct. The genome is like, an entire encyclopedia set, but a vast majority of those pages are just regulatory or structural.
3:44The XO is strictly the one to 2% of your DNA that actually codes for proteins. Ah, okay, so it's just the exact sentences that give the instructions for building the biological machines in your body. Exactly.
3:55But even though it's a tiny fraction of your DNA. It houses the vast majority of known disease causing mutations. And so make sure their findings weren't just a fluke in the UK population, the researchers replicated their work.
4:06Oh nice. Where did they look? They used Mount Sinai's highly diverse biome biobank, looking at over 28,000 participants from various ancestral backgrounds. Wow. So with this mountain of data. The researchers essentially went hunting for the Y behind disease severity.
4:21And to do it, they tested 3 specific theories using 3 distinct tools. Yes, 3 really fascinating tools. Let's start with the 1st tool, which frankly sounds like science fiction to me. It's called ESM1B.
4:35It is a 650-00000 parameter protein language model. It's incredible tech. Yeah, if you are familiar with ChatGPT. It's basically that. But instead of being trained on text from the internet, It was trained on 250000000 protein sequences across nature.
4:50What's really fascinating here is that the model is entirely unsupervised. It wasn't explicitly taught which mutations cause human diseases. The researchers didn't feed it medical textbooks or anything like that.
5:01Wait, hold on. If it doesn't know what a disease is, and it's never seen a medical textbook, how is it predicting disease severity? That sounds like magic. It's not magic. It's pattern recognition on a planetary scale.
5:14Think about how a language model works Just like ChatGPT learns that the word bark usually follows the word dog. This AI learned that certain amino acids, the building blocks of proteins always sit next to each other to fold a protein correctly.
5:28Oh I see. Yeah, it knows what a normal functioning sequence looks like across 1000000s of different organism. Ah, got it. So if mutation swaps out a single amino acid, which is what geneticists call a mis sense variant.
5:41The AI reads it like a typo. Precisely. Like if a sentence says, the dog chase the ball, and a mutation changes it to the dog chase the car, the AI says, okay, that changes the meaning, but it still makes grammatical sense.
5:53Right, it's different, but functional. But if the mutation changes it to the dog chase the Borg. The AI flags it as a severe grammatical error that completely breaks the sentence. Exactly. The AI generates a numerical score of how severely that specific typo disrupts the protein's overall function.
6:10It basically allows us to pinpoint the exact severity of the monogenic mutation itself. It's amazing. It is, but as we know, the mutation doesn't exist in a vacuum. Right, exactly. So the AI tells us exactly how broken the main gene is.
6:23But what if the rest of the person's DNA is trying to compensate for it? Or, you know, what if the rest of their DNA is secretly making it worse? How do we measure the rest of the body's genetic environment?
6:33That brings us to their 2nd tool, polygenic risk scores, or PRS. If the monogenic mutation is a sledgehammer hitting your metabolic system, the polygenic background is 1000s of people with tiny tapping hammers.
6:48Oh man, 1000s of tiny hammers. Right. PRS is a way to measure the additive background noise of 1000s of common, everyday genetic variants scattered across your entire genome. It gives us a measure of an individual's overall genetic baseline for a specific trait, completely separate from that major monogenic mutation.
7:06Okay, so we have the severity of the main mutation, and we have the additive background noise. But biology is rarely just simple addition. No it definitely isn't. If the main mutation is a lead singer singing wildly off key.
7:19The polygenic background isn't just the backup singers, it's the audio mixing board. Those background genes might actually be turning the volume down on the lead singer's mic, or they might, you know, add a distortion effect that makes the whole song sound infinitely worse.
7:33I absolutely love that analogy. And that exact concept, the mixing board interacting with the lead singer is what geneticists call epistasis. Apistasis, okay. More specifically, marginal epistasis, which is how all those background genes directly interact with the main mutation.
7:49And to measure this, the researchers had to use their 3rd tool. A novel computational method called fame, which stands for fast marginal epistis test. And from what I understand, measuring epistasis in a dataset of half a million people is notoriously difficult.
8:05Why do they need a brand new tool just for this? Because testing all pairwise genetic interactions and a biobank of 100s of 1000s of people. Well, it usually creates a computational bottleneck that literally breaks computers.
8:17Oh, wow. Just too much data. Exactly. Let's say you want to see how one single gene interacts with 20,000 other genes. Now multiply that by half a 1000000 people. The sheer number of mathematical combinations just explodes.
8:30The math gets impossible. It does. It's like trying to calculate every possible route between every single city, town, and house on Earth simultaneously. So how did fame get around the bottleneck without crashing the servers?
8:43Femi uses what is called a randomized method of moments estimator. Okay, that is a lot of syllables. Yeah. To translate the heavy math jargon. It's essentially a brilliant mathematical shortcut. Instead of calculating every single route one by one.
8:58It samples the data in a way that lets it accurately estimate the overall interaction effects. Vary. It scales linearly, meaning as the data set grows, the computing time only grows in a straight line, not exponentially.
9:10It basically bypassed the bottle mec entirely, allowing the team to conduct the 1st well powered, massive examination of marginal epistasis on disease severity. That is incredible. So we have our 3 tools.
9:22The AI protein model to judge the main mutation, the polygenic risk score to measure the background noise, and the FAMI tool to see how the background interacts with the mutation. Yep, the ultimate toolkit.
9:34Let's look at what they actually found when they pointed these tools at human data, starting with the AI model, ESM1B. I know they look closely at a gene called MC4R, which is heavily associated with monogenic obesity.
9:46Yes. And the AI model was astoundingly accurate. It didn't just flag mutations. It could successfully distinguish between fundamentally different types of mutations within that single MC4R gene. Wait, different types of mutations in the same gene.
10:01Yeah. It could separate loss of function variants, which essentially break the protein and cause severe obesity from gain of function variants. And here is the kicker. Those gain a function variants actually protect the carrier against obesity.
10:14So a disease gene isn't actually a disease gene in a vacuum. It's entirely contextual, a variant in the exact same location can either cause the disease or protect you from it. And the AI could tell the difference just by looking at the raw protein grammar.
10:29None of the older traditional variant prediction tools could do that with such precision. Exactly. a huge step forward. Okay, so what about that background noise? The polygenic risk scores? Well, to test the background noise, they set up a really fascinating comparison.
10:45They looked at people who carry a known, certified disease mutation. Then they compared them against people who don't have the mutation at all, but who simply have really bad polygenic background noise.
10:56Meaning they fall into the extreme top .one% of the PRS distribution. Right. The absolute worst case scenario for background genetics. Here's where it gets really interesting. Wait, so you're saying someone with completely normal major genes, but just a really bad roll of the dice on their background noise can actually be worse off than someone holding a certified textbook disease mutation?
11:16Yes, absolutely. For traits like high HDL and high trichlycerides, the non-carriers who simply drew a bad hand in their background genetics exhibited significantly more extreme physical phenotypes than the people carrying the established, severe clinical mutations.
11:32That's wild. We are talking about 100s to 1000s of individuals whose polygenic load results in a more extreme physical manifestation of the trait than the specific variants doctors are trained to look for.
11:44That completely flips the script on how we view a diagnosis. It means your background genetics can easily overpower the lack of a major mutation. But, um, it also works the other way, right? If you DO have the major mutation, your background genetics are still acting on you, kind of stacking up like bricks.
12:01Precisely. They proved that polygenic background has an independent additive effect on the carrier's phenotype. So it just makes everything heavier. Yep. If you carry a rarer mutation for high triglycerize and you also have a high polygenic risk score for triglycerides, those 2 things stack on top of each other, pushing your lipid levels into really dangerous uncharted territory.
12:21Which brings us to the 3rd finding, the epistosis, using that fame tool. This is where we go back to the audio mixing board. The background genes don't just add up. They multiply, they distort, they interact directly, but the main mutation.
12:35And the fam analysis revealed widespread statistical evidence of marginal epistosis with huge effect sizes. This isn't just simple addition anymore. The background variation is fundamentally modifying the effect of the monogenic variant.
12:49Oh, wow. Yeah, when they looked at the epistatic improvement percentage, which is a statistical measure of how much better we can predict the disease severity by including these complex interactions, the numbers were staggering.
13:01I have the numbers right here, and I had to read them twice. We're talking about improving predictive accuracy by up to 170% for high triglycerides. And a 48% improvement for predicting LDL cholesterol.
13:13I mean, to put that in perspective, 170% improvement means an ideal mathematical model that includes these epistatic interactions would be 2.7 times more accurate in predicting a patient's actual triglyceride levels, compared to a model that just looks at whether the person has the carrier mutation or not.
13:31That is a massive difference. It unequivocally tells us that marginal episthesis is a massive, substantial contributor to why penetrance is incomplete and why disease severity varies so wildly. This raises an important question.
13:46Well, actually, a couple of questions. How exactly are these background genes interacting with the main mutation? Are they like physically reaching over and changing the DNA sequence of the main mutation?
13:56This raises an important question indeed. No, they aren't changing the DNA sequence of the main mutation itself, but they are modifying its biological impact. And this can happen in a few fascinating ways.
14:06For instance, the background genes could be acting as regulatory switches. Like literal light switches for the gene. Exactly. They might be disrupting what we call enhancer sequences. Enhancer sequences are like hidden dials on the genome that turn gene expression up or down.
14:21Oh, okay. But even if you have a mutated gene, an enhancer might turn its production down so low that it doesn't cause much harm. Or the background genes might be causing alternative splicing. Let me guess, alternative splicing is exactly what it sounds like.
14:35Cutting and pasting the recipe differently. Spot on. It's when the body takes the same genetic instructions and splices them together in different ways to make slightly different proteins. If background genes alter the splicing of other proteins that normally interact with the mutated gene, the entire biological pathway shifts.
14:53It all connected. It is a wildly complex, beautiful web of interactions. So if we step back from the raw data and look at what this means for medicine, we are fundamentally moving from a deterministic view of DNA to a holistic one.
15:06It's no longer you have the gene, you get the disease. Precision medicine now has to integrate both rare variants and common background variants to give a patient a real, honest prognosis. And one of the most immediate, massive implications of this shift is using AI language models, like ESM1B, to reclassify what the medical field calls variants of uncertain significance or VUS.
15:30Oh, right. Let's go back to that doctor's office scenario you mentioned earlier. Right now, if you look at the Clinvar database, which is the massive public archive of human genetic variants, over 57% of misense variants are labeled as VUS.
15:43More than half. That means more than half the time a patient gets tested, and the doctor sees a single letter mutation. The official clinical report basically says, we have absolutely no idea if this is dangerous or perfectly harmless.
15:55It is an agonizing position for patients. They live with the shadow of anxiety, sometimes for years, waiting for a laboratory to grow cells in a petri dish just to see what the mutation actually does. Which takes forever.
16:05Exactly. But this study proves that ESM1B scores are tightly, reliably correlated with actual clinical phenotypes. This AI tool has the potential to reclassify 1000s of these uncertain variants mathematically, giving patients actual answers without having to wait for years of laboratory functional tests.
16:23That is a massive weight lifted off the healthcare system. But, you know, as with all massive scientific leaps, there are limitations to the study that we have to acknowledge. The researchers relied heavily on looking at electronic health records and continuous traits like cholesterol and BMI.
16:39Right, which are complicated data sources. Yeah, because we don't live in a vacuum. We live in a world with modern medicine. The researchers actually had to computationally adjust the data for patients who were taking statins trying to guess what their natural cholesterol levels would have been.
16:56Yes, and trying to find a natural baseline is a constantly moving target. Think about the current landscape of medicine, the explosive rise of modern obesity drugs, like semaglutide and other GLP1 agonists, as well as the prevalence of procedures like gastric bypass surgeries.
17:12Well, they are artificially altering natural BMI baselines on a population scale. Oh of course. As these interventions become even more common, measuring the true genetic penetrance of metabolic traits using health records, is going to become increasingly difficult.
17:26They also noted in the paper, that the study mostly focused on the longest protein coding transcript for each gene. It misses some of the deeper cellular nuances, like cell type specific protein ISO forms.
17:40Right. Which basically means the exact same gene might fold the protein into a slightly different shape, depending on whether it's operating inside a liver cell or a brain cell. The AI model in this iteration doesn't quite capture those neighborhood specific variations.
17:55And if we connect this to the bigger picture, this perfectly highlights why critical thinking is absolutely required when we look at electronic health record data. We rely heavily on EHRs for these massive bio bank studies, but the absence of a disease code in a medical record doesn't always mean the patient is completely unaffected.
18:12That's a really good point. Yeah, they might have mild symptoms. They just never bother to report to their doctor or they might be hovering right below the arbitrary clinical threshold for an official diagnosis.
18:22Right. The line between sick and healthy in a medical chart is often an arbitrary boundary we invented for insurance coding. But human biology is a continuous spectrum. Think about what happens when we eventually add our environment and lifestyle into these massive AI models.
18:36That will be the next frontier. Exactly. If our background genes act as a mixing board for disease mutation. Could our daily habits, our diet, our stress, our sleep, be the invisible hands, physically turning those dials up and down.
18:50It makes you wonder if true genomic destiny isn't written entirely in our cells, but is rather something negotiated every single day between our DNA and how we live. What does this all mean? How do we wrap up this massive shift in how we understand our own biology?
19:06It means that a genetic test is no longer a crystal ball that shows a single unchangeable future. The impact of a disease causing genetic variant isn't determined in isolation. It depends on the exact grammatical severity of the mutation.
19:20Your overall polygenic background noise, and the incredibly complex epistatic interactions between the two. By combining AI protein language models, with massive biobank scale data, we are finally moving past the outdated one gene one disease textbook model to predict individual health outcomes with true holistic genomic precision.
19:42What does this mean for the future of genetic testing when a disease gene is no longer a definitive diagnosis, but just the opening line of your genetic story? It's an opening line that we are finally truly learning how to read in context.
19:55This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
20:09If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
20:17Thanks for listening and join us next time as we explore more science based by base.