A semi-automated mtDNA reanalysis pipeline using MToolBox and MitoPhen HPO-based phenotype similarity was applied to the Solve-RD cohort, identifying previously undiagnosed mtDNA variants and adding a 0.4% diagnostic uplift.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Yeah, thanks so much for tuning in, everyone.
0:10So I want you to imagine that you're trying to put together, like, the most massive jigsaw puzzle ever made. We're talking a puzzle with 3000000000 pieces. That sounds like an absolute nightmare, honestly.
0:23Right. But you have it all laid out, and you're staring at every single piece, trying to build this complete picture of human biology. And that giant puzzle represents your standard nuclear DNA. The stuff we usually think of when we hear genetics.
0:39Exactly. But while you're obsessing over those 3000000000 pieces, there's this tiny separate bag of puzzle pieces just sitting on the corner of the table. totally ignored. And it only has what, about 16,000 pieces in it?
0:51Yeah, exactly. Just a tiny fraction. But here's the cat, those specific pieces. Power the entire picture. If they're missing or damage. The whole structure just completely fails. It's a massive blind spot in modern medicine.
1:04I mean, for 1000s of families living with rare unsolved diseases, ignoring that tiny bag of genetic material, has literally meant the difference between getting an actual diagnosis and, you know, just spending decades lost in the medical system.
1:17Today we celebrate the work of Rat Nike and the Solvard team, who have advanced our understanding of how we actually find these hidden diseases. Yeah, their 2025 paper in the American Journal of Human Genetics is just, it's groundbreaking work.
1:31It really is. Our mission for this deep dive is to look at how these researchers acted like basically genetic cold case detectives. Right, taking old evidence and looking at it in a totally new way. Exactly.
1:43They use this brilliant mix of reanalyzing old genetic data and symptom matching to finally give these families some answers. Okay, let's unpack this. Why is this specific type of DNA, mitochondrial DNA, acting like this forgotten bag of puzzle pieces?
1:57Well, to really get why it's ignored, we have to look at how genetic sequencing actually works in a clinic. When you go in for standard exome or genome sequencing, the whole system, the chemical probes, the software, it's all laser focused on the nuclear genome.
2:14Which is the 3000000000 pieces inside the nucleus. Right, but your mitochondria, which are basically the tiny power plants inside your cells, they have their own totally separate circular genome. And like you mentioned earlier, it's super small, only about 16.5 kilo bases.
2:29So because the computer's hunting for variations in this massive 3000000000 piece puzzle, what happens to the mitochondrial DNA? It just gets tossed out. The algorithms treat those mitochondrial sequencing reads as off target data, basically just background noise.
2:43Just deleted. Yeah, pretty much. But mutations in that noise are actually really common. I mean, at least one in 4,300 people has a mitochondrial disease. That's not rare at all. No, it's not. And because these diseases mess with how your cells produce energy, they absolutely devastate the organs that need the most power.
3:01So, your brain, your skeletal muscles and your heart. But the way these symptoms actually show up in real life is, frankly, a diagnostic nightmare, right? The paper talks about this specific variant, MTTL one, I think it's called.
3:15Yeah, the M.3243 AG variant. The variability there is just wild. Right. Like, you can have 2 people in the exact same family with the exact same mitochondrial mutation, and one of them just gets like diabetes and some hearing loss.
3:30Which, you know, are super common conditions. If you go to your doctor with diabetes and deafness, they are not going to jump to rare mitochondrial disorder. Exactly. But then the other family member with that exact same genetic spelling mistake develops MELAs.
3:44Right, mitopondrial encephalopathy, lactic acidosis, and stroke like episodes. It's a really severe rapid brain disease. So we're talking about the same mutation causing mild diabetes in one person and fatal strokes in another.
3:56I mean, how does the biology even allow for that? It all comes down to this really fascinating thing called heteroplasm. Hetoplasmic. Okay, break that down for us. So unlike your regular DNA, where you just have 2 copies of every gene, a single human cell can have 1000s of individual mitochondria inside it.
4:12Right. And if a mutation happens, it doesn't affect all of them at once, you end up with this mixed bag of healthy mitochondria and mutated mitochondria all floating around in the same cell. Okay, so when that cell divides to make new cells, what happens?
4:26Does it split them 50-50? Nope. It's totally random. And that is the crucial part. Over time, as your cells divide during embryonic development, this random sorting creates a bottleneck. Meaning different organs get different ratios of the bad mitochondria.
4:41Exactly. You might end up with, say, a 5% mutation load in your blood cells. So a standard blood test looks totally normal, but in your brain tissue where the energy demand is huge. That mutation load might have randomly drifted up to 80%.
4:54Which triggers the severe stroke-like episodes. Right. And because genetic testing usually just relies on a simple blood draw, these diseases are literally hiding. They're existing at these super low levels in the blood.
5:06And you'd need what, a muscle biopsy to actually find them. Yeah, something super invasive like that. Okay, here's where it gets really interesting. If this disease is actively shape shifting across different tissues and hiding from standard blood tests, how on earth do you find it in a database of 10,000 people?
5:24People who were originally tested for totally different things? Well, that computational headache was basically the whole reason the solve RD project started. It's this massive European initiative, and they gather data from over 9,900 people with rare diseases.
5:39But none of these people had answers. None. Every single person had already gone through extensive genetic sequencing, and the doctors found absolutely nothing. They were total coal cases. So instead of calling 10,000 people back in the hospital for muscle biopsies.
5:53They just looked at the old data differently. Exactly. They ran all those old files through a different software pipeline called M Toolbox. The genius of M Toolbox is that it rescues that off target data we talked about earlier.
6:05The background noise. Yeah, the discarded puzzle pieces. It takes those raw reads and specifically maps them against the standard reference map for the human mitochondrial genome. And because it's only mapping 16.5 kilobases instead of 3 billion, the computer can handle it pretty easily, right?
6:21Oh, yeah. The computational load is tiny. So they cranked up the sensitivity to catch incredibly low levels of the mutation, looking for heteroplasmy levels of just one% are higher. But wait, just finding a rare variant doesn't mean you've solved the case.
6:36I mean, humans walk around with weird, harmless genetic variations all the time. Oh, absolutely. Proving that the variant actually caused the patient's specific disease is a totally different challenge.
6:46So how did they connect those dots? That's where they brought in this really cool layer of math. They used something called human phenotype ontology or HPO terms, to actually quantify the patient's physical symptoms.
6:58So instead of a doctor just scribbling, patient has a headache on a notepad. They use structured standardized codes. Exactly. It turns messy medical notes into computable data. And then they use the database called Midafin to assign a phenotype similarity score.
7:14Okay, phener type similarity score. What does that actually look like in practice? Think of it like facial recognition software, but for diseases. Facial recognition doesn't need an exact pixel match of your face.
7:26It measures the geometric distance between specific points, like the distance between your eyes or the angle of your jaw. This algorithm does the exact same thing, but with symptoms. It measures the mathematical distance between the patient's specific symptoms and the textbook models of known mitochondrial diseases.
7:44So it spits out a probability score of how closely the patient matches the disease. I love the dating app analogy, like finding the perfect genetic match. Exactly. But I've got to push back here for a second.
7:54Isn't it incredibly risky to trust an algorithm to diagnose a life altering disease? I mean, algorithms are trained on textbook cases, but real patients are messy. Do they actually make sure this math works in the real world?
8:07That is such a valid concern, and it's why their validation phase was so important. What's fascinating here is that they didn't just fly blind. Before they touch the 10,000 unsolved cases. They tested the algorithm on a pre-solved group.
8:21Ah, okay, so a group where they already knew the answers. Right. They took 42 people with known confirmed mitochondrial diseases and hid them in a pile of over a 1000 non-mitochondrial cases. Just to see if the computer could pick them out based purely on that symptom distance score.
8:37And did it work? It did. They used the statistical tool called an ROC curve to find the perfect threshold, and they figured out that if they set the similarity score at .3, the algorithm caught 100% of the known mitochondrial cases.
8:51Wait, 100%, it didn't miss a single one. Not a single true positive. Now, a threshold of .3 is pretty sensitive. So it does flag some false positives. If they wanted to be super precise and avoid false alarms, a threshold of .48 was better.
9:04But in medicine, especially with rare diseases, that .3 threshold is literally the line between a patient getting an answer and being left in the dark. You'd rather have a false alarm that a doctor can double check than miss the diagnosis entirely.
9:18Exactly. The researchers were totally fine with a few false positives, knowing that human experts were going to manually review every flag file anyway. Okay, so they have the tool, they've proved the math works.
9:28When they finally unleashed this on the 10,000 unsolved cases. What happened? What were the numbers? Out of 10,157 data sets, the pipeline narrowed it down to 136 rare variants in 135 people. Okay. And after the chronical teams reviewed those files, it led to 37 confirmed or likely causative diagnoses, which is an extra diagnostic yield of 0.4% across the cohort.
9:54So what does this all mean? I mean, in .4% sounds like a tiny drop in the bucket. But for those 37 families. It's everything Right. These are people who went through years of hospital visits and testing, only to be told there was nothing wrong with their DNA.
10:07And now, just because someone looked at their existing data through a new lens, they finally have a name for what's happening to them. And the math held up beautifully, 92% of those 37 diagnosed patients had a similarity score above that .3 threshold, plus all the causative variants were found in their blood at heteroplasmy levels above 11%.
10:28So the disease was always there in the basic blood draw, the old computers just literally couldn't see it. Exactly. But let me ask you this. Was giving them a diagnosis just for peace of mind or did it actually change their medical treatment?
10:41Like, were any of these findings actionable? Oh, some of them were incredibly actionable. There's this one profound example from the study. They found several patients carrying a specific homoplasmic variant, M.1555 AG.
10:54Homoplasmic, meaning 100% of their mitochondria have the mutation. Right. Now, a lot of people with this mutation have no symptoms at all. They feel totally fine. But if a doctor gives them a very common type of antibiotic called an amino glycoside.
11:07Wait, just a regular antibiotic? Yeah, standard stuff. If they take it, they will suffer rapid, permanent, and irreversible hearing loss. Oh my god. A routine infection treatment causes permanent deafness.
11:19Why? It's this crazy structural quirk. That specific mutation slightly changes the shape of the patient's mitochondrial ribosomes. It actually makes them look like bacterial ribosomes. Oh I see where this is going.
11:33Yeah, the antibiotic is designed to hunt down and destroy bacterial ribosomes. So when it enters the patient's system, it gets confused, it attacks the patient's own mitochondria in their inner ear, destroying their hearing.
11:44That is terrifying. But by finding this variant in the old data, doctors can just put a big red flag in the patient's chart saying, do not give this drug. Exactly. a life-changing piece of preventative medicine all from a computer reanalysis.
11:58That is incredible. But, you know, it makes me think about the algorithm itself. connect this to the bigger picture. This whole system relies entirely on how well did doctors describe the symptoms. Yeah, that is the biggest limitation of the study.
12:10Garbage and garbage out. Right. It's like typing my car makes a noise into Google versus typing my car has a high-pitched squeal from the front left brake pad. If you aren't specific, the search engine can't help you.
12:24That's exactly what happened in the study. They actually found 3 people who definitively had severe mitochondrial mutations, but the algorithm gave them a score below .3. It said they weren't a match. Why?
12:37Because their doctors didn't write good notes. Pretty much. The clinical records had incredibly low information content. The doctors only entered like one or 2 vague HBO terms. Without specific details, the algorithm couldn't connect the dots, and the mathematical distance looked too far.
12:53So the bottleneck isn't the computer power. It's the human physician actually doing the charting. Right. And there's another huge limitation we have to talk about diversity. Let me guess. The data was mostly from one demographic.
13:0496% of the analyzed samples belong to European genetic ancestry. Oh, wow. Yeah, that's a massive training bias. It really is. Different global populations have totally different baseline mitochondrial variations that are perfectly healthy for them.
13:18If this European train algorithm looks at an African or Asian genetic sequence, it might flag a normal variation as a deadly disease, simply because it's never seen it before. So they desperately need to test this pipeline in diverse populations before rolling it out globally.
13:34Absolutely. It's a mandatory next step. Well, the core insight here is just staggering. By combining automated reanalysis of ignored DNA with highly structured symptom matching. We can literally pull diagnoses out of thin air.
13:47No new lab tests required. It completely reframes how we think about a negative genetic test. It doesn't mean the answer isn't there. It just means we didn't ask the computer the right question. This raises an important question, something for y'all to think about as we wrap up.
14:00If our clinical descriptions are actual symptoms are just as vital as our DNA in solving these mysteries. How much genetic data across the globe is currently being misinterpreted, just because we aren't describing the human experience accurately enough.
14:17And what happens when we unleash this exact symptom matching algorithm on massive mysteries, like Alzheimer's or Parkinson's disease? Could the missing genetic links already be sitting on a server somewhere, just waiting for the right map to find them?
14:30It's definitely something to mull over. It's a total paradigm shift. It really is. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
14:43If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
14:57Thanks for listening and join us next time as we explore more science, base by base.