Hundreds of genes have been reported as causes of cerebral palsy, yet there is no agreed model of what a pathogenic variant in a child with CP actually means. This study treats CP as a phenotypic feature that some genetic disorders make more likely, tests the reported genes against the population prevalence of CP across tens of thousands of published individuals, and finds statistical evidence of association for only 89 of 515, before applying that stratification to genome sequencing of 460 children with CP.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening. And don't forget to follow and rate us in your podcast app. Base by Base is now on YouTube too, at Base by Base, where every episode gets a video with chapters, and the full description, come subscribe.
0:16So imagine, if you will, a child who has just been diagnosed with the most common physical disability in childhood. Right, cerebral palsy. Exactly. And historically, the clinical assumption for CP was, well, almost entirely mechanical, right?
0:32Oh, absolutely. I mean, you looked at the chart and saw a difficult labor or maybe prematurity or some kind of hypoxic event where the brain didn't get enough oxygen. And that was kind of the end of the diagnostic road.
0:44Yeah, pretty much. Case closed. But then, you know, as next generation sequencing became faster and cheaper. The whole paradigm shifted. It really did. Suddenly the medical literature was just flooded with these rare genetic variants found in these patients.
0:56Right, we went from this purely environmental, mechanical model to this massive catalog of, I think it was 515 candidate genes that were supposedly causing this condition. Which is a huge number. It's massive But that brings us to the core of our deep dive today.
1:12What if a huge portion of those genetic answers are actually just, well, red herrings? Yeah, that's the big question. What really happens when a genetic variant found in a patient is just a statistical coincidence.
1:24Are we misdiagnosing the very foundations of this condition? We really are looking at a fundamental shakeup in diagnostic neurology here. The data strongly suggests we need to stop viewing this condition as a single distinct disease entity.
1:39Which is kind of mind blowing when you think about it. It is. Instead, it operates more as a shared phenotypic feature across a really highly specific but limited subset of genetic disorders. So it's basically forcing us to ask how many times have we sequenced a patient's genome, found some rare variant, and just incorrectly assigned it to blame just because it happened to be there.
1:59Exactly. It is guilt by association, and it's a major problem. Well, today we celebrate the work of Peter and Robinson and the team at the Jackson Laboratory for Genomic Medicine. along with researchers across the Shriners Children's Network, who have, I mean, truly advanced our understanding of the genetics behind cerebral palsy.
2:18And, you know, the clinical ecosystem here is really what makes this kind of analysis even possible in the first place. Right, because Shriners Children's operates this vast nonprofit clinical network.
2:30Yeah, and they provide long-term comprehensive care. I mean, everything from orthopedics and physical rehab to these incredibly advanced motion analysis labs. And they treat kids regardless of their ability to pay, which means they see a highly diverse population over decades.
2:45Exactly. So they have amassed this incredibly deep longitudinal phenotypic data. And in genomics today. Generating the genotype data, the actual DNA sequence is relatively easy and cheap now. Right, you just run the sample.
2:58Yeah, but getting high quality standardized phenotype data, the actual physical symptoms over time is notoriously difficult to capture accurately. It's messy. Very messy. So Shriners inadvertently created, basically, the ultimate testing ground to finally ask if these 515 genes are actually causing cerebral palsy, or if they're just along for the ride.
3:19So moving from that ecosystem to the data itself. We really have to look at how rapid this shift in CP diagnostics has been. for you and me and, you know, the medical community. It's been practically overnight in medical terms.
3:32Right. We are talking about a condition that affects roughly .3% of the global population. Which doesn't sound like a lot, but it makes it the most common childhood physical disability. Yeah, and for decades, doctors group these non-progressive motor impairments under one giant umbrella.
3:48Largely attributing them to those perinatal brain injuries we talked about. Right. But then the genomic revolution happened. Suddenly, researchers are publishing all these isolated case studies, linking CP to specific gene mutations.
4:00And that's how the literature ballooned to include 515 distinct genes. But that explosion of candidy genes must have created some serious diagnostic muddy waters, right? Oh, absolutely. Think about it from the perspective of a pediatric neurologist.
4:15They see a patient with motor impairment, and they find a rare genetic mutation. They immediately face a critical question. Is this a CP mimic? Meaning, a distinct genetic disease that just happens to look like CP?
4:29Exactly. Or is cerebral palsy a highly pleiotropic condition? Break that down for me. Pleiotropic. Basically meaning that hundreds of entirely different genetic pathways, all independently converge to cause the exact same physical motor deficits.
4:46Ah, I see. So many different broken roads leading to the exact same destination. Precisely. But until now, we haven't had a mathematically rigorous model of CP's genetic architecture to actually answer that question.
4:57Right, because to untangle a web of 515 genes, you can't just, you know, run a few dozen clinical trials, you need scale. You need massive scale. So the research has built this huge two-pronged methodology to put every single one of these candidate genes on trial, basically.
5:12Yeah, they really held their feet to the fire. First, they executed this massive literature dragnet. So they look backwards first. Right. They analyzed 21 previous diagnostic cohorts, which encompassed over 5,400 individuals with CP.
5:27And they coupled that by looking at 280 gene specific cohorts. What does that mean, exactly? Gene specific? It means they track down over 43,000 individuals in the literature, who had mutations in these specific 515 genes, regardless of whether they actually exhibited CP symptoms.
5:47Oh wow. So they're looking at people who have the mutation, but maybe don't have the disease just to see what the actual correlation is. Exactly. You have to look at the denominator, not just the numerator.
5:56Okay, so that is prong one and the second prong. The second prong was a brand new, highly controlled data set. They performed whole genome sequencing on a cohort of 460 children diagnosed with CP across the Shriners network.
6:10Right. But the sequencing was actually the straightforward part, wasn't it? Yeah, the real innovation was how they mapped the clinical phenotypes. Because they didn't just rely on humans manually reading decades of complex, highly variable doctor's notes.
6:22No, that would take lifetimes and be full of errors. They utilize an advanced text mining application called fenominal. Okay, let's dive into the mechanics of that because it's fascinating. Clinical notes are notoriously messy.
6:36Oh, they're a nightmare to standardize. Right, like one doctor might write, you know, stiff legs and another writes, hypertonia, and maybe an older chart says spastic diplegia. Exactly. And a simple control F word search can't reconcile all of that.
6:50Right. So the fenominal software doesn't just do a word search. It uses natural language processing to scan that unstructured syntax, right? Yeah, and it maps it directly to the human phenotype ontology or the HPO.
7:03The HPO, which basically translates human clinical observations into standardized computable nodes. Exactly. The HPO functions like a, well, mathematically it's a directed acyclic graph. Okay, you're going to have to translate that for me.
7:16Fair enough. Think of it as a massive, highly structured tree of medical concept. Okay, a tree. So a doctor's note that says stiff legs gets mathematically mapped up the branches of the tree to abnormality of the nervous system.
7:29Ah, I get it. It standardizes the chaos. Right. So by running fenominal, the researchers seamlessly converted 460 messy patient histories into structured digital profiles. Which they call pheno packets.
7:42Yes, pheno packets. And once they had those neat, computable packets, they fed them into a bio-informatics algorithm called Exemizer. Exomiser is so fascinating to me because it doesn't just look for broken jeans, right?
7:55It looks for biological logic. That is the perfect way to phrase it. It takes the patient's phenopacket and cross references it with vast databases of known, human, and animal disease pathways. So it's looking at mice and zebrafish data, too.
8:08Exactly. It uses network analysis, specifically these things called random walk algorithms over protein, protein interaction networks. Sounds complicated. The math is, but the question it asks is simple.
8:19If this specific gene is mutated, does it logically, biologically result in the physical symptoms we are actually seeing in this specific child? Wow, so it ranks and prioritizes the variance based on that semantic similarity?
8:32Yeah. So by standardizing the phenotype data with HPO and prioritizing the variants with Xomizer, they were able to set up a massive statistical stress test for those 515 candidate genes. Right, because the goal was to test them all against a null hypothesis.
8:50Exactly. And the null hypothesis here being that a specific gene mutation has absolutely no causal relationship with cerebral palsy, and any overlap you see is purely coincidental. Right, because remember, the baseline population rate of CP is about .3%.
9:04So statistically, you'd expect to see a few people walking around with both CP and some random genetic mutation just by pure champ. Precisely. You have to prove the association is happening more often than chance would dictate.
9:16So how do you beat that null hypothesis? The researchers employed a Poisson test for the large diagnostic cohorts. A Poisson test. Yeah, the Poisson distribution is highly effective at modeling the probability of rare events occurring within a large background over a specific interval.
9:30Right. So they used it to determine if the observed number of CP cases associated with a specific gene mutation was significantly higher than what you'd expect by pure background chance. Exactly. And then for the gene specific cohorts, they used a binomial test.
9:46Right, because they were looking at discrete outcomes for those, right? Either the patient with the mutation had CP or they didn't. Yeah, it's a binary outcome. They needed to see if the frequency of CP in that specific genetic group statistically exceeded that .3% baseline.
10:02Okay, so that is the exact mathematical hurdle. Every single one of those 515 candidate genes had to clear. And this is where the results get really, really interesting. Right, because these are 515 genes that the global medical community previously published as causing cerebral palsy.
10:18Yes. But after running them through the Poisson and binomial models, the vast majority failed. The vast majority. Yeah, the researchers found enough statistical evidence to reject the null hypothesis for only 89 genes.
10:30Wait, only 89 out of 515. That is a staggering reduction. I mean, that fundamentally challenges a huge chunk of the existing medical literature. It really does. And we see this play out so clearly in the new Shriners cohort that they sequenced.
10:45Right, the 460 kids. What happened there? Well, out of those 460 kids, they found pathogenic or likely pathogenic variants in 60 different genes. That represents a diagnostic yield of about 15.8%. Okay, so 60 gene scene problematic.
11:00But when they cross-referenced those 60 genes against their newly stress tested list of 89. Only 16 of them were actually statistically proven CP genes. Oh wow. So the other 44 were just... We were just there.
11:15This highlights a massive issue in diagnostic genomics, the bystander effect. Right. my favorite analogy for this. Just because you sequence a genome and find a mutation standing at the scene of the clinical crime doesn't mean it actually committed the crime.
11:28Exactly. It might just be a bystander watching the whole thing happen. The paper actually provides a brilliant contrast to illustrate this exact mechanism. Let's look at the gene CTN and B one. Okay, CTN and B1.
11:39This gene encodes beta catnan, which is a protein that is absolutely essential for the Wnt signaling pathway. And what does that pathway do? It regulates cell proliferation and neurodevelopment in the brain.
11:50So when this gene is mutated, the structural and developmental consequences in the central nervous system are profound. Okay, so it makes total biological sense that a mutation there would cause severe motor issues.
12:02Exactly. And the statistical models confirmed it. CTNNB1 variants are highly enriched in CP patients. It is a true causal culprit. It committed the crime. Holding the smoking gun. But then you have the other end of the spectrum with a gene like LIPH.
12:18Yes, LIPH. This is a great example. So LIPH was previously cited in the medical literature as a candidate gene for CP based on a patient case study. Right, some doctor found it and published it. But structurally, LIPH encodes an enzyme involved in lipid metabolism, and it's specifically localized to hair follicles.
12:35Wait, hair follicles. Yes. Mutations in LIPH are well documented to cause hypotrichosis 7, which is a rare hair loss disorder. So there is basically no biological mechanism connecting lipid signaling in hair follicles to the neurodevelopmental motor impairments of CP.
12:52None, zero. The patient in that historical case study likely had a mutation in LIPH, and then entirely separately suffered a perinatal hypoxic event, or maybe had an undetected mutation in a true neurodevelopmental gene that actually caused their CP.
13:08So the LIPH mutation was literally just an innocent bystander walking its dog at the crime scene. Exactly. But because it was novel, it got published and entered the diagnostic cannon as a CPG. is wild.
13:20It is a textbook example of clinical ascertainment bias. A clinician sequences of patients exome finds a rare variant they haven't seen before, and logically but completely incorrectly assumes it must be the root cause of the patient's primary severe symptom.
13:34Because they're looking for an answer, any answer. Exactly. But hold on. I have to push back a little here Rare diseases are, by definition, rare. So if we're aggressively throwing out over 400 genes just because they failed a statistical test today.
13:48Aren't we risking throwing out the baby with the bath water? I mean, maybe we just don't have enough data yet to beat the null hypothesis. Some of these variants might only exist in a few dozen families worldwide.
14:01That is a completely fair point, and the authors are keenly aware of that limitation. A lack of statistical evidence today does not equal biological evidence of absence. The failure to reject the null hypothesis for those remaining genes is heavily, heavily influenced by the limitations of current medical data.
14:19We suffer from massive publication bias in genomics. Because journals only want to publish the exciting stuff. Basically. Medical journals disproportionately publish novel or highly severe case studies completely ignoring mild or typical presentations.
14:33Which skews the underlying frequency data you need for those accurate Poissons binomial models we talked about. Exactly. It throws the math off. So it's fundamentally a structural data problem. To truly resolve the remaining 400 plus genes, we can't just rely on isolated case reports popping up in journals.
14:49We need a massive unified data set. And that requires the global adoption of interoperability standards. Right. We have to share the data. The author's point specifically to the Global Alliance for Genomics and Health, or GA4GH, and also the FHIR standard fast healthcare interoperability resources.
15:08FHIR. Yeah, that's becoming huge in Health Tech. It is. FHIR provides standardized APIs that allow electronic health records from, say, a clinic in Tokyo to seamlessly and securely share structured phenotypic data with a research hospital in Ohio.
15:23So everyone is speaking the same digital language. Exactly. We need deeply phenotyped, globally connected databases to reach the statistical power necessary to classify those ultra rare variants accurately.
15:34We just don't have it yet for all of them. So moving from the data architecture back to the actual clinic. This research demands a completely new diagnostic framework. The author's proposed what they could the phenotypic paradigm.
15:46Yes, the phenotypic paradigm. It means we have to stop categorizing cerebral palsy as a singular disease entity. Right, it's not like getting measles. No. Instead, we should define it strictly as a phenotypic feature.
15:59A symptom. Like the fever. Yes. Or the comparison to a cleft palette is highly instructive here. Oh, right. let's unpack that So a cleft palette can occur completely in isolation due to a localized developmental failure during pregnancy.
16:14Right, just a mechanical developmental hiccup. But a cleft palate is also a recognized, shared phenotypic feature across hundreds of distinct, severe mendelian genetic disorders. I see. CP operates through the exact same mechanism.
16:28It can occur in isolation due to mechanical perinatal factors like a lack of oxygen, but it is also a convergent symptom of 89 distinct mundinian diseases. So for the pediatric neurologist and more importantly, the patient and their family, this fundamentally changes the entire care pathway.
16:45It absolutely does. It paves the way for precise, actionable genomic medicine. How so? Walking me through a clinical scenario. Okay, so imagine a newborn whole genome screening reveals a pathogenic variant in a statistically proven gene like our bad guy, CTNNB1.
17:02Okay, the proven culprit. Right. The clinical team no longer has to wait and see if the child misses developmental milestones at age one or two. They know the trajectory. the math and the biology guarantee it.
17:15Exactly. So they can initiate intensive early intervention therapies before the age of two, which maximizes neuroplasticity. Because the brain is still so adaptable then. Precisely. Furthermore, if the gene points to a specific inborn error of metabolism, they can deploy targeted molecular or dietary therapies immediately, day one.
17:35That is life changing. But conversely, what if the screening finds a bystander mutation? Like our hair loss gene, LIPH. Then the clinician knows the case is still open. They haven't actually found the true etiology of the motor impairment.
17:48So they don't just say, well, we found a mutation must be CP case closed. Exactly. They must continue the diagnostic odyssey rather than closing the file on a false positive. Which prevents patients from being stranded with an incorrect genetic label.
17:59Which can have massive implications. I mean, think about family planning, recurrence risk counseling for the parents. and just accessing the correct specialized care. If you're treating the wrong disease, you're not helping the patient.
18:12Exactly. So to pull all these threads together for you listening at home, cerebral palsy is not a monolithic genetic disease. It is a complex, shared phenotypic feature, resulting from a highly specific subset of underlying Mendelian conditions.
18:28And by rigorously applying these advanced statistical models to vast clinical data sets, thanks to places like Shriners, researchers can finally filter out the genetic bystanders. Which is exactly what paves the way for true precision medicine and targeted early interventions for these children.
18:45It's incredible. It really is. And you know, it forces a much broader reckoning in diagnostic medicine as a whole. How do you mean? Well, what does this mean for how we define and study other massive heterogeneous umbrella conditions?
18:58Oh, wow. Right. Could complex diagnoses like autism spectrum disorder or epilepsy or even severe autoimmune conditions be hiding a similar landscape? A landscape of shared symptoms driven by highly distinct, statistically obscured genetic culprits?
19:12Exactly. It really makes you wonder what else we've been miscategorizing all these years. That is a fascinating thought to leave on. This episode was based on an open access article under the CCBY 4.0 license.
19:24You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast staff and leave a five-star rating. If you'd like to support our work, use the donation link in the description.
19:36Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.