Meta-analysis of up to 948,690 exome- or whole-genome-sequenced individuals across six biobanks used statistical phasing to infer compound-heterozygous genotypes, increasing detectable bi-allelic damaging genotypes by 19% and identifying 58 significant gene-trait associations, 17 of which show stronger recessive effects.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So then, I want to start by asking a question.
0:11What if the key to the next like massive blockbuster medical treatment isn't a new synthetic chemical at all? Oh right. What if it's actually finding a completely healthy person who's just walking around with broken DNA?
0:26Exactly. I mean, it sounds totally contradictory, right? Because we are so conditioned to think that genetic mutations are always the root of disease, not the cure. Yeah, we really are. But, you know, identifying people with specific inactivated genes is actually, um, it's quickly becoming one of the absolute most powerful strategies we have for discovering new therapeutics.
0:45It really is. And today we celebrate the work of the incredible research team behind a massive new meta analysis that just dropped in the American Journal of Human Genetics. We are taking a deep dive into the hidden world of what geneticists call human knockouts.
0:58Right. And the sheer scale of what they've done here is just staggering. I mean, they pooled the genetic and electronic health data of nearly a 1000000 people globally to hunt for these exact individuals.
1:09And by doing that, they are basically untangling the, um, the really complex mechanisms behind diseases and traits that affect all of us, because finding just one person who naturally lacks a functioning gene, well, it answers a huge question for pharmaceutical companies.
1:26Exactly. It tells them, hey, if we develop a drug to intentionally block this specific genetic pathway in a sick patient, is it actually going to be safe to do that? Because, you know, here's a healthy person who already lives without it.
1:38Right. They act as these naturally occurring in vivo experiments. Like, um, the paper brings up a couple of amazing real world examples right off the bat. There's this gene called PCS K 9. Oh yeah, that's a classic one.
1:50Right. Researchers found people with naturally knocked out versions of PCSK9, and these people had like exceptional low cholesterol, but otherwise they were completely fine. Yeah, they were totally healthy.
2:01And that single discovery, I mean, that biological proof of concept directly launched an entire class, a blockbuster cholesterol alluring drugs. And they also mentioned the HAO one gene, right? Finding just one healthy adult with a broken HAO1 gene gave researchers the, well, the green light to target that exact pathway for a really severe kidney condition called primary hyperoxeluria.
2:25It's just wild how one person's genome could do that. But um, to actually find these individuals on a global scale, you have to really understand the underlying mechanics of how genes work. Specifically the difference between additive and recessive genetic effects.
2:39Yeah, because a lot of complex genetics is focused on additive effects, right? Where, like, having one variant bumps your wrist up a little bit and a 2nd variant bumps it up a bit more, just sort of accumulates.
2:49But recessive effects are totally yours. It's kind of like the launch protocol on nuclear submarine. Oh, I like that analogy. Yeah. Right. To actually execute the launch command, you need 2 separate officers to turn their totally independent keys at the exact same time.
3:04And since we inherit 2 copies of every gene, you know, one from mom, one from dad, a true knockout requires a recessive effect. Exactly. Both copies of the gene have to be broken. Because if only one key is turned, the submarine doesn't launch.
3:19The remaining functional copy just picks up the slack and the body produces enough of the protein to function totally normally. So finding the people where both keys are turned is, um, apparently a massive computational hurdle, I guess the easiest scenario to spot is when a person inherits the exact same mutation on both copies.
3:38which we call being homozygous. The variant on the maternal chromosome is perfectly identical to the variant on the paternal chromosome. It's a clean match. But the paper points out that there's a 2nd much stealthier way to get a knockout, right?
3:51being compound heterozygus. Yeah that's where it gets really tricky. That's when both copies are broken, but the mutations themselves are totally different from each other. Like, maybe the mom's copy has a premature stop code on and the dad's copy has a, um, a frame shift mutation or something.
4:07Exactly. Both of those airs destroy the resulting protein, but they do it in completely different ways. And identifying those compound heterozygous variants. That is where the math gets incredibly complex.
4:20Okay, but wait, why is that? Because if you run my genome through a sequencer and it flags 2 severe mutations in the exact same gene, shouldn't the computer just be able to say, oh, both copies are broken, it's a knockout?
4:34Well, it would be that simple if standard sequencing, read your entire chromosome continuously from end to end. But, you know, it doesn't do that at all. Oh, because it chops it up. Right? Right. Standard short read sequencing, chops your DNA into 1000000s of tiny little fragments, reads them, and then tries to map them back to a reference genome.
4:53So when the algorithm spots a mutation over here and another mutation nearby, it completely loses the structural context. So it literally can't tell if both mutations are sitting on the exact same chromosome, which what is that called?
5:05That's called being in cis configuration. Right, insists. Or if they are on opposite chromosomes, which would be in trans. Exactly. And that distinction is everything. Because if both mutations are insist like.
5:17They're both just sitting on the maternal chromosome. Well, the paternal chromosome is still perfectly intact. Ah, so the submarine doesn't launch. The gene still works. You only have a true knockout if the mutations are in trans.
5:29You nailed it. And since the sequencer jumbles all the fragments, you have to perform this process called phasing, basically reconstructing the parentage of the individual chromosomes. And historically, wouldn't you just have to sequence the person's parents to figure that out?
5:43You would. But doing that on a massive scale is impossible. So the researchers here use this really advanced technique called statistical phasing. Instead of needing the parents' DNA, they use massive reference populations and algorithms.
5:58How does that actually work, though? Like mathematically. Well, the algorithms look for these established patterns of genetic inheritance. They're known as hapotype blocks within global populations. And by comparing the fragmented sequencing data against those known patterns, they use probability to infer which mutation is sitting on which chromosome.
6:16Wow. So they are statistically guessing the structure and it works. Because the paper says by applying this across 100s of 1000s of people, they increase their discovery of these true compound knockouts by an impressive 19%.
6:28Yeah, a 19% bump is huge in this field. But, um, to actually train an algorithm to do that, you need an unprecedented amount of data. Which they definitely had. They pulled together a staggering collaborative data set.
6:41We're talking 948,690 individuals. Almost a 1000000 genomes. It's incredible. Across 6 totally distinct global biobanks, like the UK Biobank, the all of us research program here in the US, Genomics, England, Biobank, Japan, and they evaluated all that against 41 distinct clinical traits.
7:01And what's really fascinating is just the sheer volume of discovery that that kind of scale enables. They ended up identifying 5,563 total gene knockouts. That's a lot of broken DNA. It is. It actually expanded the known universe of entirely knocked out human genes by nearly 20%.
7:19They found 1767 totally new genes that we didn't even know could be safely inactivated in humans. And we have to talk about the demographics of those newly discovered genes because that is vital. Out of those 1767 new knockouts, 1371 of them were found in individuals of non-European ancestries.
7:41Right, particularly within South Asian subcohorts in the data. Yeah, and it really just exposes this massive blind spot in historical genetic research. I mean, studying only one population, which historically has almost always been European data.
7:54It's like, I don't know, trying to catalog all the rare birds on earth by exclusively walking through a forest in Germany. That is a great way to put it. You're never going to find the species native to the Amazon or the Arctic, if you don't look there.
8:04Different populations have totally different historical bottlenecks and migration patterns. Plus, the paper mentions that in some cultures, there are higher rates of endogamy and consanguinity, right? Yeah, absolutely.
8:15When parents share a recent common ancestor. It dramatically increases the rate of auto-sygosity. That's where identical chromosomal segments are inherited from both sides. Which naturally drives up the occurrence of those homozygous knockouts we were talking about earlier.
8:29Precisely. It makes sequencing these globally diverse populations mathematically essential. If you don't do it, the statistical power to find novel drug targets just isn't there. So, okay. They assemble this incredibly diverse database.
8:45They find over 5,500 broken genes. The next logical step is seeing what actually happens to the health of the people carrying them. Right. They cross-referenced all those genetic knockouts directly with the individual's electronic health records, and they found 58 significant associations between a broken gene and a specific clinical trait.
9:04And they were really rigorous about making sure these were true recessive effects, right? Very. Out of the 58, they isolated 17 distinct instances where the clinical impact was strictly driven by that 2 key knockout mechanism, not just some additive effect masquerading in the data.
9:20And looking at those 17 associations, the intersection of the genetics and the hospital records is just deeply revealing, let's talk about the PYGMG. Oh, this is one of my favorite findings in the whole paper.
9:30It's so wild. Because when both copies of PYGM are completely knocked out. The data shows a massive statistical association with elevated AST levels. And AST is an enzyme, right? So if you go to a hospital for routine blood work and your AST is spiking, the standard medical assumption is that your liver is damaged.
9:50Right. Doctors look at AST and immediately think hepatic issue. But the thing is, PYGM is famously a muscle gene. It has practically nothing to do with the liver. Wait, so if it's a muscle gene, why is it causing a liver enzyme to spike in the hospital records?
10:03Because of the fundamental limitations of how we co diagnoses. The knockout of the PYGM gene actually causes something called macardal disease, which is a really rare recessive glycogen storage disorder.
10:15Okay. So because the muscles can't properly break down glycogen for energy, physical exertion literally damages the muscle cells. They undergo lysis, they basically break open. Ah, and when those muscle cells rupture.
10:27They release all their AST directly into the bloodstream. So the patient shows up at the clinic, the blood panel flags, high AST, and the hurried physician just codes it into the electronic health record as a potential liver problem.
10:40Wow. So the symptom is captured totally accurately in the record, but the actual biological mechanism is completely misidentified. Exactly. And they observe the exact same phenomenon with a gene called OD 81.
10:52In the health records, people with OD 81 knockouts frequently carried a diagnosis of COPD chronic obstructive pulmonary disease, which, you know, is incredibly common. But I'm guessing an OD 81 mutation doesn't actually cause standard like smoking induced COPE.
11:08Nope, it causes primary ciliary dyskinesia. The microscopic hairlike structures in the respiratory, track the cilia, they failed to beat properly, so they can't sweep mucus and pathogens out of the lungs.
11:19Which leads to chronic airway obstruction, which clinically looks almost identical to standard COPD. Right. And if we connect this to the bigger picture, it just fundamentally changes how we've used standard medical databases.
11:32What this proves is that a significant number of quote unquote common diseases, clogging up hospital records, are actually rare Mandelian genetic disorders just hiding in plain sight. That is mind blowing.
11:45They are miscategorized because our diagnostic tools are built around grouping symptoms together rather than identifying the root genetic cause. Yeah, and moving from a symptom-based diagnostic model to a mechanism-based one is going to reveal some incredible biological paradoxes.
12:00A single genetic mechanism influencing multiple, totally contradictory trait. Which perfectly sets up the mystery they found with the HBB gene. Oh, the HBB paradox. is fascinating. It really is. So the HBB gene is arguably one of the most well-known genes in all of genetics.
12:16Mutations there are the primary cause of severe red blood cell disorders, like sickle cell disease and beta thellasemia. But when this meta analysis looked at people with knocked out HBB genes specifically in the African and South Asian cohorts, a totally counterintuitive cardiovascular profile emerged.
12:35Very counterintuitive. Right. These individuals exhibited significantly lower cholesterol levels and a lower body mass index. Which, if you just look at those metrics in complete isolation, you might assume the knockout provides some robust protective benefit for your heart.
12:48Except they also displayed a dramatically higher risk of heart failure. Now, my media thought reading that was, while aren't the lipid and heart issues simply secondary side effects of suffering from a really severe chronic illness?
13:00The very logical assumption. Right, because if a patient is battling severe hereditary anemia, their physiological stress is immense. Like malnourishment could cause BMI and cholesterol to drop and chronic hypopsia could just overwork the heart until it fails.
13:16Isn't the algorithm just picking up the systemic fallout of a known disease? The researchers anticipated that exact confounder. So to isolate the direct effect of the gene from the secondary effects of the illness, they performed what's called a conditional analysis.
13:31What does that mean in this context? It means they statistically removed every single individual who carried a clinical diagnosis of sickle cell disease, or beta thalasemia, from the data set. I just erased them from the pool to see if the cardiovascular associations would vanish.
13:45But the association's held firm. Even without the diagnosed disease present in the data, the HBB knockouts still strongly correlated with lower cholesterol and increased heart failure. Exactly. The genetic mechanism itself is driving a very distinct cardio metabolic profile.
14:00The elevated risk of heart failure is likely tied to iron overload cardiomyopathy. Because the red blood cell turnover is compromised. Yeah. Iron starts accumulating in the myocardial tissue of the heart.
14:12eventually causing it to fail. And at the exact same time, the body detects the compromised red blood cells, and aggressively attempts to synthesize new cell membranes. Oh I see. Yeah, that rapid, desperate synthesis acts as a massive metabolic sync.
14:27It pulls huge amounts of available cholesterol right out of the blood plasma to build those membranes. Yeah. So the patient's overall cholesterol score absolutely plummets. That is incredible. What does it say to you that a single genetic tweet can protect your cholesterol while simultaneously threatening your heart with iron overload?
14:44I mean, it completely shatters the binary concept of mutation being strictly beneficial or strictly harmful? It really highlights the intense pleotropy of our genome. Yeah, you know, where one gene regulates multiple, totally disparate systems.
14:57There's rarely a biological free lunch. Altering a fundamental pathway, almost always demands a trade-off. And we actually see that exact same contradiction when we look at physical traits in the study.
15:08Like, breaking a gene doesn't always lead to a deficit. The paper analyzed human height and found 2 fascinating associations. Yeah, the height data was super interesting. First, they found that knocking out a gene called LECT2 leads to a decrease in height.
15:23Okay, makes sense. But then they found an entirely uncharacterized gene. It doesn't even have a formal name yet, just the identifier. E-N-S-G 000000267561. Catchy name. very catchy. When this specific gene is knocked out, It actually increases height.
15:41Which is highly unusual. Finding a knockout that enhances a complex polygenic trait like height is really rare. extremely rare, especially when you consider the broader rules of population genetics. The paper explicitly notes that autozygosity, you know, inheriting identical broken copies from related parents, almost universally decreases human height across the board.
16:01Right, it's a classic manifestation of inbreeding depression, a general reduction in overall biological fitness. So how does breaking this random, unnamed gene manage to do the exact opposite and make people taller?
16:13Well, it really comes down to the highly localized function of the proteins involved. Even though this specific gene is uncharacterized, it's situated near genomic regions that we know regulate height.
16:24And there is some strong evidence pointing toward its involvement in selenium metabolism. Selenium, how does that affect height? Well, selenium plays a really critical role in regulating oxidative stress and thyroid hormone metabolism within chondorcytes.
16:39Condrocytes or the cartilage cells, right? Exactly. They're the specialized cells that produce and maintain cartilage. So if you knock out a gene regulating that specific metabolic pathway, you alter the cellular environment of the cartilage.
16:52Precisely. And the growth plates at the ends of our long bones are made of this exact cartilage. Normally, as we age, these plates ossify and fuse, which is what halts our vertical growth. Oh, wow. So if disrupting the selenium pathway alters the hormonal signals within those cells.
17:07It could potentially delay the fusion of the growth plates. A slight delay in fusion gives the bones more time to elongate, resulting in a taller individual, even though the broader biological system might actually be experiencing inbreeding depression.
17:22That entirely defies the general rule, which really underscores why empirical base-by-base mapping is so vital. We just cannot rely on phenotypic assumptions anymore. We really can't. So, to synthesize all this for your listening.
17:36We started out discussing how difficult and murky genetic diagnosis can be. But the central insight here is that being well informed in this space today means recognizing that the methodologies are finally catching up to the complexity of our biology.
17:49Absolutely. By combining massive, ancestrally diverse global databases, with incredibly clever statistical phasing, researchers can finally cut through the noise. They can separate the cyst from the transmutations, locate the compound heterozygotes, and identify the true human knockouts.
18:06And this computational heavy lifting is literally the only way to find the specific biological targets that will lead to the next blockbuster treatment that changes 1000000s of lives. What does this mean for the future of clinical practice in your eyes?
18:19I think the overarching implication of the Lassen paper is that global collaboration is the absolute prerequisite for the future of medicine. The genetic architecture of human biology is just too intricately woven for any single biobank or any single ancestral population to unravel on its own.
18:37It requires a unified international effort. It really does. And that leaves me with a final somewhat provocative thought for you to explore on your own. If a meta-analysis of nearly a 1000000 genomes proves that standard diseases in our hospital records things like common COPD or routine liver enzyme spikes are frequently just rare recessive genetic disorders misdiagnosed by outdated coding, what happens to medicine when we sequence a 1000000000 people?
19:01a huge question. Yeah, when massive statistical phasing connects a 1000000000 genomes to a 1000000000 health records will the very concept of a common disease cease to exist altogether? We may soon realize there is no such thing as standard asthma or standard heart failure, and that the entire medical dictionary will be replaced by 1000s of highly personalized distinct genetic fingerprints.
19:24It's going to be a completely different world. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
19:36If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode and inspired by the article you've just heard about.
19:50Thanks for listening and join us next time as we explore more science base by base.