The Undiagnosed Diseases Network (UDN) applied joint whole‑genome analysis across 4,236 individuals and introduced RaMeDiES, an analytical framework to prioritize genes by de novo recurrence and compound heterozygosity while integrating intronic splice predictions and experimental validation. The work recapitulated known diagnoses, identified new diagnoses and candidates, and released software and a public browser to enable cross‑cohort, deidentified discovery.
0:19Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. glad to be here for this one. Yeah, so imagine for a 2nd that you are sick.
0:33But not just like a regular kind of sick. Right, something entirely different. Exactly. You have a disease so rare, so incredibly unique that the absolute best doctors in the world have just never seen it before.
0:46Yeah, which is terrifying. It is. You are essentially pushed onto what the medical community calls a diagnostic odyssey. You go from specialist to specialist. You endure these endless tests, scans, blood draws.
1:00And nothing comes of it. Right. Every single time, the results come back totally inconclusive. You are alone, clinically speaking, just stranded on an island with this mystery illness. It's a devastating place for a patient to be.
1:11Truly. But now, what if I told you that analyzing the DNA of thousands of other people, people with, you know, completely unrelated, ultra rare diseases, could finally solve your specific mystery? It sounds counterintuitive.
1:23I mean, looking at unrelated diseases to solve yours. Yeah, but today, in this deep dive, we are shifting our perspective entirely. We are putting down the clinical magnifying glass that just looks at patients one by one.
1:38Right, moving away from that isolated view. Exactly. And instead, we are looking at a massive crowd of diverse patients all at once, using some incredibly advanced math. So what really happens when we let algorithms hunt for the mathematical shadows of disease across a whole population at once.
1:56Today, we celebrate the work of the undiagnosed diseases network, alongside researchers at Harvard Medical School and collaborating institutions who have advanced our understanding of rare genetic diagnosis.
2:07Yeah, and this group is just doing phenomenal work. They really are. And, you know, to understand the gravity of this work, you kind of have to realize that this network acts essentially as the court of last resort for medical mysteries in the United States.
2:18Right, like the final stop. Exactly. By the time a patient is even accepted into this specific cohort, they have, well, they've completely exhausted every traditional medical avenue. Wow, yeah. We are talking about individuals suffering from severe neurological regressions, completely unexplained immune failures or, you know, complex cardiac anomalies that just defy textbook explanations.
2:43So they really are the toughest of the tough cases. Yeah. Traditional medicine has hit an absolute wall, and their doctors are just completely stumped. Okay, let's unpack this because historically, finding the genetic cause of the disease relied on a very different set of tools, right?
3:01Right. a much more narrow approach. Yeah, if you were a geneticist a couple of decades ago, you basically had 2 primary methods. You could track a specific physical trait through this massive multigenerational family tree.
3:14Like tracing it through a pedigree, yeah Exactly. Or you could group together a large bunch of patients who all had the exact same highly uniform symptoms. Which is great if you have a lot of similar patients.
3:25Right, but I always think of it like trying to find a typo in an instruction manual for how to build a human body. Oh, that's a good way to put it. Thanks. Like, the traditional way is to find a bunch of people reading the exact same manual.
3:36Because machines are breaking down in the exact same way. And they all point to the same misspelled word on, you know, page 42. Right. The homogeneity makes it obvious. Yeah, but with this cohort of patients, it's like you are the only person who has ever received your specific manual.
3:52Exactly. Your machine is breaking down, but nobody else in the room shares your symptoms. So how do you find your typo if you can't compare your broken parts to anyone else's? Well, the answer requires a fundamental paradigm shift in how we actually handle human genetics.
4:08We are uh, moving deeply into a data science phase now. moving away from the human eye, basically. Yeah, exactly. For a long time, diagnosing these ultra rare, isolated cases relied heavily on individual clinical intuition.
4:23Right, doctor just guessing based on experience. Pretty much. A doctor looks at a single patient's genome, looks at their symptoms, and just tries to manually deduce which of the 1000s of genetic variants might be the culprit.
4:34Which sounds exhausting. It is. Human intuition is incredibly powerful, but it has severe limitations when you're tasked with sifting through the 3 billion base pairs of the human genome. Yeah, the math just doesn't favor the human brain there.
4:49Exactly. So the field is now moving toward joint genomic analysis. Rather than looking at one patient in isolation. Researchers are taking phenotypically broad cohorts. And in groups of people with just wildly different physical symptoms, right?
5:02Right. And analyzing their genetics collectively. The goal here is to establish rigorous statistical proof of disease causing genes, basically moving away from biological guesswork. And to do that effectively, you must need a staggering amount of data.
5:16Oh, absolutely. The scale is massive. Yeah, because we are looking at harmonized whole genome sequencing across 4236 individuals for this analysis. And crucially, that number includes 1463 complete family trios, meaning researchers sequence the child suffering from the disease, along with both of their unaffected parents.
5:38Which is key. Yeah, that isn't just an ocean of genetic data. It's a really incredibly structured data set. The architecture of that data is vital. And the technology they developed to navigate it is a software suite called ramedies.
5:50Yeah, which stands for Rare Mendelian Disease Enrichment Statistics. And, you know, what sets ramedies apart is its strict genotype 1st approach. Wait, a genotype 1st approach means the algorithm completely ignores the patient's physical symptoms at the beginning of the analysis.
6:08That's right. It doesn't look at the symptoms at all, initially. Why would you blind the software to the very clinical presentation you're trying to cure? I mean, isn't the whole point to figure out why the patient is sick?
6:20I know, it seems totally counterintuitive at 1st glance. But if you tell the algorithm what the symptoms are right off the bat, you risk introducing human bias into the search. Oh, like a self-fulfilling prophecy.
6:31Exactly. You might accidentally restrict the algorithm to only look at genes we already associate with those specific symptoms. Ah, so you'd miss anything completely novel. Precisely. By blinding the software to the clinical presentation, the algorithm looks purely at the statistical probability of the genetics.
6:48Just the raw math. Just the math. It asks a purely mathematical question. Is this specific genetic mutation happening more often across this entire diverse group of 4000 patients, then random chance would allow?
7:01Wow. Okay, so how does it actually test that? Well, the software suite is broken down into specific statistical tests. The 1st is called remedies DN, which hunts for de novo mutations. DeMovo meaning new.
7:15Like, these are mutations that the parents just don't possess at all. Exactly. They spontaneously appear in the child's DNA during early development. Finding those sounds pretty straightforward since they stand out from their parents' DNA.
7:26But what makes this new software better at finding them? Well, the older methods for finding these rely on what geneticists call permutation models. And those models are computationally just exhausting.
7:37Why? What do they have to do? They basically had to randomly scramble the genomic data 1000s and 1000s of times to simulate a baseline of what random chance actually looks like. Just to see if one mutation was significant.
7:49Exactly, just to see if an observed mutation was statistically significant. But Romody's DN abandons that cumbersome method entirely. Thank goodness Yeah. It utilizes an advanced base pair resolution mutation rate model called roulette.
8:04Roulette. like that name. Yeah, and this model already knows the inherent baseline probability of a mutation occurring at literally any single letter in the genome. Oh, so it skips the scrambling entirely.
8:15Right. It then pairs that natural baseline rate with cutting edge AI pathogenicity predictors. Specifically tools like alphemisense and splice AI. Okay, I've heard of those. Yeah. They predict how damaging a given mutation is likely to be to the final protein.
8:31Okay, let me make sure I understand the AI component here because it's so easy to just say AI and assume it's like magic. Right. It's not magic. It's very good modeling. So alphemisance isn't just acting like a basic spell checker looking for typos in the DNA.
8:46It's functioning more like a structural engineer reviewing a blueprint, right? That is a great analogy, yeah. Thanks. Like, it calculates whether swapping out one specific molecular brick in a protein will cause the entire biological bridge to collapse, right?
9:01That captures the mechanism perfectly. The software takes the natural speed limit of mutations, like how often this typo happens just by chance, and multiplies it by the AI structural engineers assessment of how badly that typo will break the resulting protein.
9:17That's incredible. It is. And using those 2 metrics, it calculates a mutational target score. It does this fast, I assume. Oh, incredibly fast. Because this approach relies on a direct analytical equation rather than scrambling data 1000s of times, it can calculate these probability scores across 1000s of genomes in mere seconds.
9:35That is a massive upgrade. Yeah, but the software's capabilities extend beyond just spontaneous mutations. It also tackles a much more difficult genetic scenario called compound heterozygous variants. And it does this using a tool called ramedy's DCH.
9:49Okay, compound heterozygotes. This is when a child inherits two different broken copies of the exact same gene. Like one broken copy from the mother and a different broken copy from the father. Exactly.
10:02Two hits to the same gene. I know finding these mathematically is notoriously difficult. Primarily because of something called population structure, right? Yes, population structure is a huge headache for geneticists.
10:13Because human populations aren't perfectly mixed like a deck of cards. People who share geographic or cultural ancestry tend to share rare genetic variants just by historical chance, completely unrelated to any disease.
10:27Yeah, so population structure actually completely confounds traditional statistical math and genetics. I mean, if an algorithm ignores how human populations are, you know, geographically structured, it just fails.
10:39How so? It will look at your cohort and flag a bunch of inherited rare variants as statistically significant, simply because several people in your study happen to share ancestry from the same specific region of the world.
10:50Oh, I see. So it thinks it found a disease gene, but it just found a regional marker. Exactly. It creates massive amounts of false positives. This is where I'm a bit stuck on the mechanics of the solution.
11:01The literature says ramedy's bypasses this trap by using the parents' genomes as a localized baseline. Right. But if my parents share my geographic ancestry, wouldn't their genomes also be biased by that same population structure?
11:16How does looking at the parents actually fix the map? It fixes the math because the algorithm completely changes the question it is asking. Okay. It doesn't try to map out the entire historical population structure of every single patient in the massive cohort that would be impossible.
11:33Instead, it conditions the math on the specific rare variants that were actually passed down in that specific family trio. Oh, yeah, I'm falling. Yeah, so the algorithm looks at the mother's rare variants and the father's rare variants.
11:45And it asks, given the pool of rare variants, this specific mother and father possess, what are the mathematical odds that they both passed down a highly damaging mutation that landed in the exact same gene in their child?
11:58It focuses entirely on the probability of 2 independent rare transmission events colliding in the exact same biological locations. Oh, wow. By narrowing the focus to the transmission odds within the trio, the overarching population structure becomes totally irrelevant to this specific mathematical test.
12:16You nailed it. It completely removes the background noise. So it's looking for the astronomical unlikelihood of 2 independent genetic lightning strikes hitting the exact same gene in the child, based solely on what the parents had hovering in their clouds.
12:30That is exactly what it's doing. That makes a lot of sense. Here's where it gets really interesting. Weren't older tests ignoring the deep non-coding parts of our DNA? How is this deep dive into the genome capturing the dark matter of our DNA that older exome sequencing missed?
12:45This really highlights the critical advantage of whole genome sequencing over older methods. Historically, diagnostic testing largely relied on exome sequencing. Because it was cheaper and easier. Yeah, much easier.
12:57But the XOM is just the one to 2% of our DNA that actually codes directly for proteins. We largely ignored the other 98%. The vast stretches of non-coding DNA called introns. Exactly. We ignored them because we just lacked the computational tools to interpret them reliably.
13:14But remedies brings deep intronic variants into the exact same statistical framework as the coding variants. Which is revolutionary. It relies on AI tools like splice AI, to predict if a mutation hidden deep within the dark matter of an intron, will disrupt the way the gene is ultimately spliced together and translated.
13:34And when you look at the results of running this advanced math over those 4000 genomes, the numbers are incredibly validating. They really are. The analysis successfully recapitulated over 80 known diagnoses.
13:46So right out of the gate. We have proof that the math works, because it successfully found the needles we already knew were hiding in the haystack. Exactly. It proved its own accuracy first. But beyond that, it established five entirely new diagnoses and three new putative ones.
14:01Let's dig into some of these discoveries, because this is where the power of the diverse cohort really shines. The De Novo breakthroughs are particularly striking. A prime example is the LRRC 7 gene. The algorithm flag 2 different patients in the cohort.
14:14Now, these individuals did not have the exact same overarching disease presentation. Right, they weren't carbon copies of each other. Not at all, but they shared overlapping symptoms, specifically hypotonia, which is low muscle tone, and severe developmental delays.
14:30And the math revealed that both of these completely unrelated patients had spontaneous mis sense mutations in a very specific microscopic region of the LRRC 7 gene, known as the Lucine Rich Repeat region.
14:45And without analyzing a massive, diverse crowd simultaneously, those 2 isolated cases with their subtle overlap might have remained medically unsolved forever. Almost certainly. Another compelling example is the HRC 5 genes.
14:58This is a histone gene. Okay, let's get some context there. For sure. His stones are the essential protein spools that our DNA thread wraps around to stay organized inside the cell nucleus. Very important stuff.
15:08Crucial. The algorithm flag 2 patients with spontaneous mutations in this gene, both presenting with infantile onset motor delays and distinct facial dysmorphologies. And what's wild is that older tests completely missed this.
15:22Yeah, the critical detail here is that this specific gene was not flagged as dangerous by older, simpler analytical tests. Because it didn't look broken enough. Exactly. The integration of alphemisense, our AI structural engineer, was required to recognize that these extremely rare mutations would fundamentally destabilize that crucial protein school.
15:42But the most fascinating finding for me centers on that dark matter we discussed earlier. Oh, the Emmy 11 case. Yes. There was a specific patient in the cohort, suffering from severe neurodegeneration and Korea.
15:53Which manifests as these involuntary, unpredictable muscle movements. Yeah. And the answer to their diagnostic odyssey was completely invisible to traditional XO testing. What's fascinating here is how well hidden it was.
16:06Both of the inherited disease causing mutations for this patient were hidden deep in the first intron of a gene called Emmy the 11. So totally outside the coding region. Exactly. If a clinician only ordered standard XM sequencing, the Exxons, the protein coding regions would appear perfectly healthy.
16:23It was only by utilizing whole genome sequencing that the researchers could even see these deep entronic variants, ramedys flag them because the AI predicted they would trigger a disruptive biological process called cryptic splicing.
16:36Okay, cryptic splicing, let's break that down, because it's central to why non-coding DNA matters. It's great concept to visualize. Yeah, if we imagine the cell's genetic machinery as a film editor putting together a movie, the Exxons are the crucial scenes that make it into the final theatrical cut.
16:53Right, the good footage. And the introns are the outtakes, the unused footage left on the cutting room floor. A cryptic splice mutation, essentially tricks the editor. It does. It creates a false cut here signal in the dark matter, tricking the cellular machinery into leaving a massive chunk of garbage footage in the middle of the final movie.
17:11Exactly. The protein becomes unwatchable, or in biological terms, non-functional. That is a brilliant way to explain it. And to confirm that this cryptic splicing was actually occurring in the ML 11 gene, the researchers took a step beyond the DNA math.
17:26They needed physical proof. Right. They performed RNA sequencing using the patient's blood. Because the RNA acts as a transcript of the final edited movie. Precisely. And the RNA sequencing explicitly proved that the cellular machinery was, in fact, retaining that first intron.
17:43The splicing mechanism was completely broken by that mutation in the dark matter. So we have the math finding the broken code, the AI predicting the mechanism of the failure, and the RNA sequencing, providing the physical proof.
17:56It is a stunning convergence of technologies. It really is. But the researchers didn't stop at linking single genes to diseases, right? They also took a wider view to look for broken biological pathways.
18:07Yes, which is the next frontier. How did they connect these isolated genetic typos back to the actual physical symptoms the patients were suffering from? Well, biological functions rarely rely on a single gene operating in isolation?
18:20They rely on entire pathways? Like an assembly line. Exactly. Intricate cascades of multiple genes working in concert? The researchers recognize this, and cluster the patients not by their specific genetic mutations, but by the similarity of their physical symptoms.
18:34How do you mathematically cluster a symptom? They utilize the human phenotype ontology, which is essentially a highly standardized hierarchical dictionary of human symptoms. Oh, so it allows computers to group patients logically based on how their diseases physically manifest?
18:49Right. One of the clusters they generated was a group of 15 patients with varying neurological symptoms. Okay. And when the algorithm evaluated this specific group, it didn't find one single broken gene shared among them.
19:03So no single smoking gun. No. Instead, it found that these patients were mathematically enriched for mutations in a specific biological pathway involved in taste transductions. Now I have to stop you here.
19:14Because I think anyone listening to this deep dive would be incredibly confused. It sounds bizarre, I know. Yeah, why would a pathway responsible for helping you taste your food, be driving severe debilitating neurological disease.
19:26It seems entirely disconnected until you look at the fundamental bavology of how a taste bud actually works. Case transduction relies heavily on specialized ion channels and receptors to send rapid electrical signals to the brain.
19:39Electricity. Yeah. We are talking about mechanisms like calcium channels and gabba receptors. Within this neurological cluster of patients, the algorithm found mutations in different distinct genes, specifically CAC N1C and Gaber 3.
19:56And what do those genes do? Both of these genes are responsible for building those rapid fire electrical signaling channels. Oh, I see. So a taste bud is essentially just a specialized sensory neuron. Exactly.
20:07If you have a broken ion channel in your tongue. You might just lose your sense of taste. But if that exact same broken ion channel pathway is operating in the central nervous system, in the brain itself, the resulting electrical misfiring causes catastrophic neurological symptoms.
20:21That is the crucial insight. It proves that sometimes the answer to a rare disease isn't a single shared gene, but a shared functional pathway that operates across different tissues in the body. That is wild.
20:33It is. Identifying these shared pathways is a massive leap forward because it opens the door to umbrella therapeutics. Meaning one drug could treat multiple different genetic diseases. Exactly. The possibility of treating different individual genetic flaws with a single drug designed to stabilize the entire shared pathway.
20:52But to perform that kind of sweeping pathway analysis, or to confidently identify deep intronic mutations, you need an immense amount of data. Oh you need oceans of it. And that brings us to a significant systemic bottleneck in global genetics.
21:08Patient privacy. Yeah, international privacy laws are incredibly strict regarding medical data and for very good reason. A hospital in London cannot simply email the raw, fully sequenced genome of a patient to a clinic in Boston to run joint statistical analyses.
21:25Which makes sense. You can't just email a patient's entire biological blueprint around. Consequently, genomic data remains heavily siloed across different institutions and countries. I'm stuck on something here regarding the solution, though.
21:36The literature states that ramedy solves this by generating summary level mutational target statistics. This means a hospital doesn't share the patient's actual DNA sequence. They just share the mathematical probability score of the mutations they observed.
21:50Exactly. Just the math, not the DNA. But if hospital A only sends a summary score. And hospital B only has a summary score, how do they know they are looking at the exact same broken gene if neither of them is allowed to see the actual DNA letters?
22:05It's actually really clever. They know they're looking at the same gene because the summary statistics are indexed to specific genomic coordinates. Oh, like a map. Exactly. Think of the genome as a vast map with specific GPS coordinates for every single gene.
22:19Okay, I'm with you. Hospital A runs the ramedy's algorithm locally, entirely behind their own strict firewalls. The algorithm looks at the gene located at a specific coordinate and calculates the aggregated mathematical probability that the mutations found there are driving disease in their local patients.
22:36And then what? Hospital A, then exports only that mathematical risk score tied to that specific coordinate. They share that anonymous privacy safe mathematical summary with researchers in the U.S. So no patient data leaves the building.
22:49None. The U.S. researchers can then align that coordinate score with the score from their own local patients at that exact same coordinate. They are combining the mathematical weight of the evidence without ever meeting to look at the raw underlying DNA sequence that generated it.
23:05So what does this all mean? We can finally share the mathematical shadow of the genome to find matches without compromising a single patient's privacy. That is exactly what it means. And as a proof of concept, the researchers actually did this.
23:19Really? Yeah, they securely cross-analyze data from massive international databases, like the deciphering developmental disorder study, Genadex, and the Radbood University Medical Center. Wow. It completely proves that global collaboration is mathematically possible without compromising a single patient's privacy.
23:38That is incredible. But, you know, to maintain scientific rigor here. We do have to acknowledge that this current approach is not flawless. No, of course not. It has clear limitations. What's the main hurdle right now?
23:50Well, the Romdy's model is incredibly precise for point mutations, like single letter typos and small insertions or deletions. However, it currently struggles to analyze structural variants. Structural variants being like massive missing chunks or large scale duplications of DNA.
24:08Right. Right. It also does not accurately model the untranslated regions at the very edges of genes. Why does it struggle there? The reason for this blind spot is that the scientific community simply does not yet possess accurate baseline mutation rate models for those massive, complex genetic events.
24:25We just don't have the mafia. Right. To overcome this, the field will need to transition away from current short read sequencing technologies toward long read sequencing. Which can map out those massive structural changes with much greater clarity.
24:38Exactly. It's the necessary next step for the technology. Looking at the entire landscape we've covered today, the central insight is this. The era of diagnosing rare disease is purely on a case-by-case basis, relying solely on human clinical intuition, has really reached its limit.
24:54It absolutely has. But by shifting to joint, statistically calibrated analysis, across massive phenotypically diverse cohorts, algorithms can uncover the intricate patterns that human eyes simply cannot see.
25:07Right, the math reveals what we can't. Yeah. By leveraging whole genome sequencing to illuminate the dark matter of our DNA, and by sharing privacy safe mathematical summaries, the global medical community can finally collaborate to solve the most difficult genetic mysteries on Earth.
25:24We are truly witnessing a transition from clinical isolation to global mathematical collaboration. What does this mean for the 1000000s of people worldwide, still waiting at the end of their own diagnostic odysseys?
25:36What if the cure for the rarest disease on Earth is hidden in the data of someone with a completely different illness on the other side of the planet? It's a profound thought. This episode was based on an open access article under the CCBY 4.0 license.
25:49You can find a direct link to the paper and a license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating. If you'd like to support our work, use the donation link in the description.
26:03Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about. Thanks for listening and join us next time as we explore more science, base by base.
26:31On nights with bright screens and patient files, a thousand quiet puzzles, a thousand miles, We trace the letters Where the small brakes hide, looking for a name to stand beside. Not just one story. Not one lone skin.
26:52We stitched the cohort like a living map in our hands from base pair whispers to a splice gone wrong. We turn the noise into a signal strong. Run it together. Let the patterns ignite. Deidentify light.
27:11De identify life. Do you know what goes in the to L.A. accord? Point to the doorway we were missing before? Make the hidden feel right. They identify light. They identify like. Deep in the entrance, where the warning won't show up, distance, which can then the transcripts flow, one assay, and any variants on the line.
27:54We watch the splicing shoes, it's designed. Calibrated chances steadying clear. Not a miracle, just truth drawing near. New diagnosis arising from the haze and candidates that hold their gaze run it together.
28:13Let the patterns ignite, de-identify light. They identify life. From denying white oceans to a single heart, cool, fine white exams Could never get through. Running together. Make the hidden feel right, de identify right.
28:34Deidentify light.