Baya et al. applied a misalignment framework to UK Biobank polygenic scores and exomes and found that individuals whose observed phenotypes deviate from polygenic expectation are enriched for rare damaging variants across multiple traits and diseases.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So imagine for a 2nd that you decide to take one of those, you know, highly detailed, comprehensive genetic tests.
0:15Right. The ones that look at 1000000s of genetic markers. Exactly. So you get the results back and the report boldly predicts, based on all those common genetic markers, that you should be exceptionally tall.
0:26But you read this completely perplexed because, well, in reality, you are quite short. Yeah, or consider a scenario that carries a lot more weight. Say the genetic test predicts you have an extremely high risk for severe heart disease.
0:39Oh, wow, yeah. But year after year, your doctor confirms your arteries are just perfectly clear. So the big question is, what happens when your actual physical traits completely defy your genetic destiny?
0:51It's fascinating, really. I mean, we spend a tremendous amount of computational power building these massive predictive models in modern genomics. Yet the extreme outliers, the individuals who just simply refuse to fit the mold, they're often brushed aside a statistical noise.
1:07Which is a shame. It really is. Treating them as anomalies is a huge missed opportunity because their defiance of genetic expectations is actually one of the clearest, most powerful biological signals we have for understanding how human biology truly works.
1:23Right. Think of your overall genetic background, like a massive weather forecast. You have 1000000s of common genetic variants coming together, creating a general prediction based on broad patterns. Like a high chance of rain or a heat wave.
1:35Exactly. Maybe the forecast indicates a high risk for a certain disease. Yeah. But suddenly what happens when a highly localized, incredibly rare genetic lightning strike just completely overrides that entire forecast.
1:47That lightning strike is everything. It is. So today, our mission in this deep dive is to explore the science of phenotibic misalignment. We are going to look at what causes people to deviate from their genetically expected traits.
2:01And crucially, how studying these genetic rebels can unlock entirely new treatments for rare diseases. Today we celebrate the work of Nicholas Baia, Duncan Palmer, and their co-authors from the University of Oxford, the Broad Institute, and their associated institutions.
2:17Yeah, they have really advanced our understanding of complex disease architecture here. They created this powerful, systematic framework to study the interplay between our common genetic background and those ultra rare, high impact variations.
2:31Okay, let's unpack this, because to really appreciate the magnitude of what this research team accomplished, we need to understand the baseline they were working from, right? And I guess that starts with understanding how we currently predict genetic destiny using polygenic scores or PGS.
2:45Right, so a polygenic score is essentially a mathematical way to quantify your genetic risk for a disease or a trait, and it's based entirely on common genetic variants. And we all carry 1000000s of these, right?
2:55We do. And individually, a single common variant has almost no impact. Like, it might increase your risk of heart disease by just a fraction of a percent. Barely noticeable on its own. Exactly. But when you aggregate 1000000s of them together into a single score, well, they function as a highly effective tool for stratifying a population's baseline disease risk.
3:16I like to think of this in terms of a personal financial budget. Your common variance, so your polygenic score, represent your steady, predictable baseline. Yeah, your fixed monthly income. Right, your income and your standard recurring expenses.
3:29It provides a very reliable forecast of your financial health over time. But the blind spot of a polygenic score is that it entirely ignores the rare variance. And that blind spot is exactly what this new framework addresses.
3:42Because polygenic scores are built strictly on common variants, they just cannot account for those rare, high impact mutations. Which brings us to the liability threshold model, right? Yes, a really foundational concept in genetics.
3:54Developing a complex disease is not just a simple light switch that flips on or off. Instead, it's a continuous scale of accumulating liability. So what makes up that liability? It's a combination of things.
4:06Your baseline, common genetic variance, any rare genetic variance you might happen to carry, and of course your environmental exposures. So picture a cup sitting under a dripping faucet. Right. The water level rising in the cup is your accumulating liability.
4:21You only actually develop the disease when that total liability crosses a specific threshold. Basically causing the cup to finally overflow. Exactly. So returning to our budget analogy, the liability threshold model makes perfect sense.
4:33Your polygenic score is your stable baseline budget. But a rare, highly damaging genetic variant acts like a devastating, unexpected medical bill that completely ruins your overall financial status in an instant.
4:46Regardless of how stable your monthly budget was. Right. Or on the flip side, a rare protective variant acts like a sudden massive lottery win. It shields you from financial ruin, even if your baseline income is just terrible.
4:59That's a perfect way to put it. So the central challenge then is figuring out how to actually find those rare genetic lottery wins and unexpected bills in a massive C of data. And preventing disease relies on accurate modeling.
5:13It does. If we can understand why diagnosed individuals sometimes have an unexpectedly low common variant risk, we can uncover those hidden monogenic, meaning single gene, drivers of disease. So to do this, the research team must have needed just a staggering amount of data.
5:29Oh, absolutely. They used the UK biobank focusing on over 400,000 individuals with European genetic ancestry. Wow, 400,000. Yeah, and they didn't just look at standard markers. They analyzed over 25000000 specific variants derived from 450,000 exome sequences.
5:47For little context, while whole genome sequencing looks at all your DNA, XM sequencing specifically isolates the regions of the DNA that actually code for proteins. Right, it's focusing purely on the functional machinery of the body.
5:59Which is exactly where those rare high impact mutations usually hide out. So what traits were they actually looking at with all this data? They targeted 2 specific types. First, they looked at 7 continuous traits.
6:10These are things that exist on a spectrum. Like height and weight. Yeah, standing height, BMI, bone mineral density, and LDL cholesterol. And second, they looked at 3 dichotomous traits or case control traits.
6:23Binary conditions, essentially. You either have them or you don't. Exactly. Type 2 diabetes, coronary artery disease and osteoporosis. But to find these outliers. I mean, they couldn't just look at absolute numbers, right?
6:35They had to compare the genetic forecast against the real world. Right. The misalignment classification. This is the critical step. How does that actually work? Well, 1st they adjust the observed real world traits for basic covariates, like age and sex.
6:48So an 80 year old isn't unfairly compared to a 20 year old. Precisely. Then they compare that adjusted real-world number to the genetically expected phenotype, which is the polygenic score. Okay, so matching expectations versus defying them.
7:00Right. They grouped people into aligned, meaning their real traits match their genetics and misaligned, meaning they hit those higher, low extremes. And once they isolated these genetic rebels, the misaligned group, they ran them through a really rigorous testing funnel to hunt for rare variants, right?
7:19Yes, a 3 stage funnel. Stage one focused on canonical genes. These are the usual suspects, genes we already know are linked to known monogenic disorders. stage two. Stage 2 broadened the search using the GEL panel app genes.
7:34That's a larger expert curated list of diagnostic grade genes for rare disorders. Okay, and then the final stage. Stage 3 was a completely unbiased XM wide scam. They just searched across the entire Xome for entirely new discoveries.
7:49But across all 3 of those stages. They had to define what damaging actually means, right? Because we all carry 1000s of mutations. Yeah, but most of them are biologically silent. They don't do anything.
7:59So the team used strict prediction algorithms to look for 2 specific types of damage. was the 1st one? Predicted loss of function variants or PLOFFs. This essentially breaks the gene entirely. You can think of it like tearing entire pages out of an instruction manual.
8:13So the protein just doesn't get made at all. Exactly. And the 2nd type was damaging misense variants. These alter the protein structure. So it might not destroy it completely, but it severely impairs its function.
8:27Like a bad typo in the instruction manual. Yes, exactly. Okay, so what does this all mean for the data? If someone deviates from their expected cholesterol, couldn't it just be because they are taking a statin or eating poorly rather than some rare genetic mutation?
8:43That is a great point. I mean, the environment seems like the easiest explanation for misalignment. It absolutely is. And the researchers knew this. They actually checked for this very thing. They found that non-genetic factors, like socioeconomic status, smoking, and yes, taking cholesterol alluring medications did heavily associate with misalignment.
9:02So there was a ton of noise in the data. A tremendous amount, but that makes the genetic discoveries they did make even more impressive. Oh I see. Yeah, the environmental noise made the true genetic signal incredibly hard to find.
9:13So the fact that they found these rare variants cutting straight through the noise of statins and lifestyle choices. It proves just how robust the genetic signal really is. The rare variant really is that blaring air horn drowning out the entire symphony.
9:27Exactly. So let's look at the actual findings. Starting with the continuous traits and those canonical genes, the usual suspects. The statistics here are just staggering. Let's take height. People who were significantly shorter than their common genetics predicted were highly enriched for those predicted loss of function variants in genes like A can, IGF one and SHOX.
9:49Aikin specifically stood out, right? It did. The odds ratio for the ACAN gene was a massive 367. An oz ratio of 367. I mean, in human biology, researchers are usually thrilled to find a variant with an odds ratio of like one.
10:022.2. Oh, definitely. An oz ratio of 367 means that if you possess a broken variant in this gene, your mathematical likelihood of being exceptionally short is 367 times higher than someone without it. It's an absolute biological sledgehammer.
10:16And it makes sense biologically. The ACAN gene provides instructions for cartilage structure. If you break it, your bone scaffolding is compromised, halting vertical growth no matter what your other genes say.
10:28And what about people who are taller than expected? They were enriched for damaging misense variants in the FBN one gene, which is linked to connective tissue elasticity. Wow. And what do they find regarding cholesterol?
10:41Well, people with lower than expected LDL cholesterol had a combined burden of broken variants in the PCS canine and APOB genes. Which are deeply involved in how the body clears cholesterol from the blood.
10:53Exactly. And they also looked at bone mineral density or BMD. This is a trait that had never really been studied this way before. Oh, really? What did they find there? People with lower than expected BMD were highly enriched for mutations in the KopiB2 and Gorab genes.
11:07And those are linked to a rare form of osteoporosis, aren't they? Yes. So by just looking for people whose bones were weaker than predicted, they successfully flushed out the individuals carrying these rare disease mutations.
11:19Here's where it gets really interesting. How does this framework apply to complex diseases? You know, conditions you either have or don't have, like type 2 diabetes? The case control results are fascinating.
11:31They completely prove the liability threshold model. We talked about earlier. So? Let's look at type 2 diabetes. The researchers found patients who definitely have diabetes, but who also possess a very low polygenic risk score.
11:45So based on their common genes, their liability cup should be nearly empty. They shouldn't be sick. Right. But they were significantly enriched for rare, highly pathogenic variants in MODY genes, specifically HNF1A and HNF4A.
12:02MODY stands for maturity onset diabetes of the own, right? That's right These genes control the beta cells in the pancreas that secrete insulin. So if you get a severe mutation there, your baseline genetic risk doesn't matter.
12:13Your rare variant just overwhelmed your good, common genetics. That is the ultimate unexpected medical bill ruining the budget. Exactly. But what about the other side? The unexpected lottery wins. Did they find people whose baseline budget was terrible, but they were perfectly healthy?
12:27They did in coronary artery disease or CAD. They found healthy control subjects, people with perfectly clear arteries, who possessed a dangerously high polygenic risk score. So they absolutely should have had her heart disease.
12:40Yes, but they were enriched for rare protective variants in the AG PTL 3 gene. So essentially, their rare variants acted as a shield, protecting them from their own bad common genetics. Exactly. It's an incredible protective mechanism.
12:54And then they moved on to the 3rd stage, right? The unbiased exome wide scan. The total unknowns. Yes. The scan identified 74 significant genes across the genome. What stood out there? Well, for BMI, they found that lower than expected BMI was tightly linked to the NPL gene.
13:12And NPL is super interesting because its biological function spans across species, right? Yes, the paper notes that a deficiency in this specific gene causes muscle loss not just in humans, but in zebrafish and mice as well.
13:25Which is wild, seeing that mechanism conserved across such vastly different species gives us so much confidence that this is a true biological function. Absolutely. They also linked lower than expected BMI to the ACSL 6 gene, which is involved in lipid synthesis.
13:41Okay, and what about the other traits in the XOM scan? They looked at age at menopause, too. Higher than expected age at menopause was linked to damaging mis sense variants in the CanK1 team. Okay. Yeah, it offers some really tantalizing clues about extrogen dependent gene expression in the reproductive system.
13:58So connecting all of this to the real world. We have this massive data set. We've identified the outliers. We found these rare variants. What are the actual implications for clinical practice? The most immediate application is triage.
14:10This misalignment framework proves that polygenic risk scores aren't just for predicting common diseases. They can actually be used as a powerful triage tool to identify which patients should be prioritized for expensive, rare variant genetic screening.
14:25Because you can't just sequence everyone's X summits too expensive. Exactly. But if a patient is sick and their polygenic score says they shouldn't be, That misalignment is a massive red flag pointing toward a rare monogenic driver.
14:36And beyond triaging patients. I imagine this acts as a sort of treasure map for drug development. Oh, absolutely. It fundamentally changes target discovery. Think about the ACSL 6 gene we just mentioned.
14:47The one linked to lower BMI. Yes. Damaging that gene results in a lower than expected BMI because it impairs lipid synthesis. So, theoretically, pharmaceutical companies could design an antibody to specifically target and inhibit ACSL 6.
15:03Basically, using a drug to artificially recreate the exact protective effect of that rare genetic mutation to treat severe obesity. Precisely. That is incredible. But we are talking about finding needles in a genetic haystack here.
15:17What are the limitations? What are the blind spots of this methodology? There are a few important caveats. The biggest one is simply statistical power. Because they're looking for ultra rare variants in a tiny, restricted group of misaligned people, the raw numbers are minuscule.
15:32How small are we talking? In many of the tests, literally 0 people in the misaligned group happen to carry the variant. It makes it mathematically very hard to prove associations. Yeah, double rarity leads to incredibly small sample sizes.
15:45Another limitation is that this is largely computational. Right, algorithms rather than doctors. Exactly. The variants were categorized as damaging by computer prediction algorithms, not by manual critical review of the patients.
15:58And while the algorithms are good, they aren't perfect. And finally, there has to be a limitation regarding global diversity, right? Because polygenic scores require so much baseline data. Yes, a major limitation.
16:10This study was restricted exclusively to individuals with European genetic ancestry. Because the background linkage patterns of DNA differ between populations. Right. Creating accurate polygenic scores requires massive ancestry specific data pools, and those are currently lacking for non-European populations.
16:29So until we have diverse global data, we can't easily apply this misalignment classification to everyone. Well, if we distill this entire study down, what is the central insight here? Ultimately, the core insight is that phenotypic misalignment. When a patient's real world traits defy their common variant genetic predictions, is a powerful biological signal.
16:50We can't just ignore the outliers anymore. No, we have to investigate them. By studying these genetic rebels, We can validate the complex liability of disease, improve diagnostic screening for rare monogenic disorders, and discover entirely new therapeutic targets.
17:05What does this mean for the future of personalized medicine? When your doctor looks at your chart, will they treat the genes you have or the genetic expectations you've managed to defy? This episode was based on an open access article under the CCBY 4.0 license.
17:21You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
17:33Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.