Boston Children's Hospital used VS-NN, HPO NLP, and DRAGEN reprocessing in a proactive genomic reanalysis to identify candidate diagnoses in 2% of pediatric ES/GS cases.
0:00Welcome to Base by Base, the papercast that brings Genonics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So for today's deep dive, I want you to imagine something.
0:11Imagine you take a medical test to find out what's causing a severe illness. Right. Maybe it's for you, or maybe it's for your child. Exactly. You wait weeks, you're incredibly anxious, and finally, the results come back, but they just say inconclusive.
0:24Ugh. The worst possible answer. It is. So you put that piece of paper in a drawer, you try to move on, and you just sort of live with the mystery. But what if that piece of paper could automatically update itself the exact moment, science finally discovers the answer?
0:39That is the dream, right there? It really is. And it's the premise of this massive shift happening in genomic medicine right now. Today we celebrate the work of Boston Children's Hospital, who have advanced our understanding of something called proactive genomic reanalysis.
0:54Yeah, PGR. It's um, it's essentially the death of the static medical document. Right. And the birth of this continuous dynamic diagnostic engine. So, okay, let's unpack this. Why is this Boston children's study so revolutionary?
1:11What's the blind spot? they're trying to fix? Well, to really grasp it, You have to look at the current status quo. Over the last decade, whole exome and genome sequencing. They've become the absolute gold standard for diagnosing rare diseases.
1:24Sure, they're incredibly powerful tools. They are, but despite all that computational power, a staggering 55 to 67% of patients with a presumed genetic condition, they actually remain undiagnosed after their 1st test.
1:38Wait, really? More than half? Yeah, more than half. So you take the most advanced genomic tests available? And more than half the time, you still walk away with 0 answers. That is, I mean, that's wild.
1:48It's incredibly frustrating for families, and the reason comes down to the friction between the speed of science and the reality of clinical practice. Science is moving at just a breakneck pace. Right, like new discoveries every week.
1:59Yeah, every single day, actually. New gene disease links are published globally all the time. A genetic variant that was classified as, say, a variant of uncertain significance back in 2019. Dreaded V US.
2:13Exactly. That VUS might be definitively linked to a specific neurological disorder by 2024, but it's not just the science that evolves. It's the patient. Oh, because especially in pediatrics, the clinical picture is a moving target.
2:27Spot on. The phenotype, the actual physical symptoms, they change as the child grows. So a two-year-old might present with a pretty generic developmental delay, which, you know, doesn't narrow the genetic search down much at all.
2:40Right. It could be a 1000 different things. But then by age 5, they might develop a highly specific focal seizure pattern. So the symptoms come into focus. Science is understanding comes into focus. But that original genetic report is still just sitting in the drawer.
2:53Exactly as it was printed years ago. It's, um, it's like buying a smartphone that never ever gets a software update. If medical science is constantly updating its operating system, Why is our personal genetic report frozen in the exact year we took the test?
3:07It shouldn't be. And the frustrating part is that the medical establishment actually knows this is a flaw. They do. Oh, yeah. The accepted best practice is that a patient's genomic data should be reevaluated periodically.
3:19Okay, so why isn't it happening? Because the current mechanism requires the treating clinician to manually order a reanalysis from the laboratory. And the baseline data on how often this actually happens is just bleak.
3:33How bleak are we talking? The Boston Children's Study found that only 13.6% of patients ever received this clinician initiated reanalysis. Wow. So almost 87% of people just never get a 2nd look at their data.
3:47Why is that bottleneck so tight? Mostly, it's an infrastructure failure. Healthcare data is heavily siloed. The clinical data, you know, the new seizure patterns, the latest specialist checkups that lives in the hospital's electronic health records, the EHR.
4:01Right, the doctor's computer system. But the raw genomic data, the actual DNA sequences, those usually live on the proprietary servers of third-party reference laboratories. So the hospital has the clinical puzzle pieces, and the lab has the genomic puzzle pieces.
4:14Exactly. We're talking about massive outside companies like Genadex, or legacy labs like Cleritus Genomics, and there is 0 automated crosstalk between them. And I imagine these reference labs aren't just going to run complex, computationally expensive reanalyses for free unless a doctor specifically demands it.
4:34Nope, which puts this massive, unfair burden on an already overworked physician, or, frankly, on the most proactive vocal parents to constantly call the lab and ask for updates. It's a static system that just fundamentally doesn't scale.
4:47You can't rely on individual human memory to constantly cross reference 1000s of patients against a globally evolving database. You really can't. Which brings us to the engine Boston children's build, because if human memory doesn't scale, you have to centralize and automate the process.
5:02And they took a pretty massive swing here, right? Huge. They brought the raw genomic data of over 15,500 patients and their relatives back in-house. So they essentially liberated this data from the external lab servers.
5:15Yes, and integrated it directly into the hospitals' own platform. That centralization was the necessary 1st step. Because once you have the data securely in-house. You can build an automated pipeline to continuously analyze it.
5:29Right. So they developed an AI driven system. It features a variant scoring neural network or VSNN. VSN, got it. But the genius of this system isn't just the genetic analysis part, it's how it bridges the gap with the clinical data.
5:44Because the system has to know how the patient is doing today, not 5 years ago. Exactly. To do that, the system uses natural language processing or NLP. Like the tech behind AI chat bots. Very similar.
5:56On a monthly basis, this NLP software scans through the patient's clinical notes in the electronic health record. Wow, it reads everything. It reads the physician's shorthand, the specialist reports, the updated symptoms, and automatically extract those evolving clinical features and converts them into standardized medical codes.
6:13Are these those HPO terms? Yes, human phenotype ontology terms. That introduces an interesting vulnerability, though. I mean, clinical notes are notoriously messy. Oh, they are a disaster sometimes. Right.
6:25Dr. A writes. Muscle weakness, Dr. B writes hypotonia. Dr. C writes some completely obscure medical shorthand. If the NLP misinterprets that messy text. Doesn't it just feed junk data into the neural network?
6:38That's a great point, and that's exactly why the HPO translation is so critical. The NLP is specifically trained on clinical vernacular to act as a universal translator. Oh, okay. It recognizes that muscle weakness and hypotonia actually map to the exact same standardized HPO code.
6:56It takes the subjective chaos of human medical notes and structures it into a rigid, machine readable vocabulary. Okay, so the computer now has a constantly updated standardized list of the patient's current symptoms.
7:08What does the neural network actually do with that? Well, the algorithm evaluates the patient's genetic variants against 43 different features? 43 like what? Things like the frequency of the variant in the general population.
7:20Because if a mutation is common, it's likely harmless. It cross-references the variant against those newly updated HPO symptom tags to see if there's a clinical match. And crucially, it runs the DNA through advanced AI prediction tools, specifically models like alphemous sense and splice AI.
7:38Okay, let's double click on those. Because throwing around names like alphamus ends sounds incredibly impressive. But how are these tools actually predicting if a mutation is dangerous? Let's break down the mechanism.
7:50Alphamous and specifically looks at mis sense mutations. That's where a single letter change in the DNA results in the substitution of a single amino acid in a protein. Just one amino acid gets swapped.
8:02Right. And before AI, guessing if that one swapped amino acid would break the whole protein was incredibly difficult. Alpha my sense uses deep learning based on structural biology to model how that protein folds.
8:14So it visualizes the shape. Exactly. It essentially predicts whether that specific amino acid swap will structurally collapse the protein, or if it won't really make a difference. It's sort of like replacing a single Lego brick in the middle of a massive Lego castle.
8:27Alphemescence runs a physics simulation to see if swapping a standard brick for a slanted brick causes the entire load bedding wall to collapse. That's a perfect analogy. And splice AI does something similar, but for RNA splicing.
8:40How does that work? When DNA is transcribed into RNA, the non-coding sections have to be precisely cut out and the coding sections spliced together. Right. editing the transcript. Splice AI predicts whether a mutation will trick the cellular machinery into cutting the RNA at the wrong spot, which would scramble the resulting protein entirely.
8:59Oh, wow. So the neural network takes all of this. The alpha isn't structural predictions. The splice AI predictions, the population frequency, the HPO clinical matches calculates all 43 features, and then what?
9:11It spits out a pathogenicity score between 0 and one. Okay, I understand the mechanism now, but I still have to push back here on behalf of the listener. Go for it. An algorithm calculating a score between 0 and one is not a physician.
9:22Having a neural network effectively diagnose a rare genetic disorder sounds a bit risky. Are we taking the human doctor out of the loop? Not at all. And the researchers were hyper aware of the risks of algorithmic hallucination or, you know, overcalling pathogenic variants.
9:40So humans are still in charge? Completely. This is strictly a semi-automated system. Think of the AI as an incredibly aggressive spam filter. It doesn't permanently delete the harmless variance, but it routes the highly suspicious ones directly to a human reviewer's primary inbox.
9:56What's the threshold for that? Any variant scoring .5 or higher triggers a mandatory 2 pass human review. Okay, so the AI acts as the filter, not the final authority. It raises its hand and says, hey, human experts, look at this one.
10:09How rigorous is that human review? Highly rigorous, but highly efficient. Because the AI isolated the target. First, a trained bioinformatician reviews the variant. And how long does that take? They spend about 5 minutes doing a technical check.
10:22They look at the raw sequence data to ensure the read depth is sufficient. Basically making sure the sequencing machine didn't just make a typo. Right, a quality check. And they check the inheritance patterns to see if it makes genetic sense.
10:36If it passes that hurdle, it goes to the 2nd pass. Who does the 2nd pass? A licensed genetic counselor. They spend about 3 minutes doing a deep dive into the clinical criteria, verifying the exact genotype phenotype correlation in the established scientific literature.
10:52Five minutes in 3 minutes. That is lightning fast. It is, but it's only possible because the AI cleared away the massive haystack. The human experts don't have to wade through tens of thousands of harmless mutations.
11:04They just evaluate the handful of needles the AI found. It maximizes the value of human clinical expertise by removing all that computational busy work. But, you know, building a massive AI safety net is one thing, exposing it to the messy reality of undiagnosed patients is another.
11:21When they actually turn this machine on and scan the backlog, did the safety net hold, let's look at the pilot steady results. Yeah, the pilot was fascinating. They ran this algorithm on 62,637 genetic variants from 2144 patients who had absolutely no previous diagnosis.
11:40We're talking over 2000 medical mysteries. And the funnel worked exactly as designed. Out of those 62,000 variants, the AI flag just 310 is highly suspicious. So it threw out almost everything. It did.
11:53The Human Review team then analyzed those 300 in, and they narrowed it down further to 45 variants in 42 patients. And then those 42 files were sent to the actual treating clinicians at the hospital. Right.
12:04And the confirmation rate is stunning. Clinicians confirmed that for 33 patients, these variants had a high suspicion of disease causality. Plus another 3 were identified as strong candidates requiring just a bit more clinical follow-up, right?
12:16Exactly. 33 definitive diagnoses retrieved from the dark. So what does this all mean? Why were these missed initially? That is the core scientific value of this pilot? The breakdown perfectly validates why proactive, continuous reanalysis is so necessary?
12:30Of those confirmed cases, 13 were found because brand new gene disease relationships were published in the scientific literature after the patient's original test. So the data was always there. Science just hadn't caught up to it yet.
12:43Right. Another 3 were found because the original testing laboratory had actual technical limitations in their bioinformatics pipeline that missed the variant entirely. But BCH's new, more sophisticated pipeline caught it.
12:57Exactly. And crucially, 2 were found specifically because the patient's physical symptoms had evolved over time to perfectly match the gene profile. Oh, so the NLP caught the phenotypic expansion in the clinical notes, fed it to the network, and the match was made.
13:12You got it. Think about the gravity of those numbers. 33 families who have been living in diagnostic limbo, going to specialist after specialist, suddenly getting a definitive answer, not because they endured another invasive biopsy or a new blood test, but simply because an algorithm looked at old data with a new lens.
13:30It is the ultimate promise of precision medicine realized, but, you know, this brings us to the most sobering part of the study. The messy human element of healthcare. Exactly. The transition from a computer finding a variant to a patient actually getting a diagnosis in the real world, because the bioinformatics pipeline was a massive, elegant success.
13:50But the real world of healthcare logistics, that proved to be incredibly chaotic. Yeah, finding the mutation in the code is really only step one. The logistics of delivering that answer revealed massive systemic cracks.
14:03Let's talk about the clearius genomics situation because that perfectly illustrates the vulnerability of our current data architecture. It is a prime example. Many of these unsolved cases came from sequencing done between 2015 and 2017 by an external lab called Clearedus Genomics.
14:18The Claritus went out of business. They close down. Their servers went offline. Which means the genomic data was effectively orphaned. Exactly. Because the original reference lab no longer existed, Boston Children's couldn't follow standard protocol.
14:32What is this standard protocol? Usually just ask the original lab to run a free digital reanalysis to confirm the AI's findings. Instead, BCH had to scramble to find internal pilot funds to literally pay for new physical confirmation tests at entirely different laboratories.
14:49Wow. The infrastructure to seamlessly support these legacy patients simply didn't exist once the commercial entity vanished. It's like having your entire digital life locked inside a software platform that suddenly goes bankrupt, and you have no way to export your own files.
15:05And the administrative friction went far beyond bankrupt labs. There's the issue of provider turnover. Oh, right. In 24 of these newly solved cases. The doctor who ordered the original genetic test years ago had left the hospital.
15:17Think about the administrative cascade that triggers. The hospital had to identify a brand new specialist, assign them the case, and get them up to speed on a complex, undiagnosed patient they had never actually met.
15:30Just so a physician could legally evaluate this new genetic finding and deliver it to the family. Right. And that leads to the ultimate bottleneck, patient recontact. Because we are talking about time gaps of five, sometimes 8 years.
15:43Families move across the country. Kids age out of pediatric systems, transition to adult care, and just get entirely lost to follow up. The study mentions some families had relocated internationally, making care coordination a total nightmare.
15:57And tragically, the researchers noted that 2 of the patients whose data finally yielded an answer had already passed away before the reanalysis was complete. That is just heartbreaking. It highlights a painful truth about genomic medicine.
16:10We have spent 1000000000s developing the sequencing technology and the artificial intelligence to interpret it, but we have severely underivested in the human infrastructure required to actually maintain longitudinal relationships with these families.
16:23It is. It's like discovering you hold a winning lottery ticket from 5 years ago. But the bank that issued it has dissolved. The manager left the industry, the person whose name is on the ticket moved overseas.
16:35And you have to somehow track them down just to hand them the prize. Reading through the methodology you realize the neural networks, the structural biology models, the NLP, that was actually the easy part.
16:47The administrative friction of the healthcare system was the true bottleneck. And that bottleneck has profound implications for medical equity. How so? Well, in the traditional manual system who gets a reanalysis, it's the families with the resources, the time and the health literacy to constantly advocate for themselves.
17:06The squeaky wheels. Exactly. What Boston Children's has demonstrated is a technological solution to an equity problem. A centralized, automated background system means every single patient's data gets reevaluated continuously, fairly, and systematically.
17:21Regardless of whether their parents have the time to call the clinic every 6 months. Right, it levels the playing field. The algorithm doesn't care about your socioeconomic status. It just runs the code.
17:30But the stark reality the researchers point out is that to make this model work on a global scale, hospitals cannot just license the ANI. They have to secure dedicated funding, permanent infrastructure, and specialized personnel whose entire job is logistical follow-up, and patient recontact.
17:48The AI can find the needle, but humans still have to sew the thread. That is a crucial takeaway. The technology is practically magic, but it's useless that the final mile of delivery is broken. So to bring all these threads together.
18:02Boston Children's Hospital has unequivocally proven that genomic data should never be a static document. By combining advanced AI filtering leveraging tools like NLP and structural protein modeling with hyperefficient human expertise, hospitals can actively mine legacy DNA data to uncover life-changing diagnoses for rare diseases.
18:21The model works. It works beautifully. And it paves the way for a near future where your genetic code acts as a living, breathing entity within the medical system. Like a continuous subscription service.
18:31Exactly. Your sequence genome sits on a secure, centralized server, constantly scanning global research databases while you sleep, cross referencing against new variant discoveries and your own evolving electronic health record.
18:44The moment a new scientific consensus matches your biological reality, the system flags it. But that leaves you with one final thing to consider, something to mull over. If algorithms are continuously scanning our genomes in the background, updating our files as science evolves, the very definition of a diagnosis changes, it's no longer a noun, it becomes a verb.
19:07It's a continuous process. Right. And while that is medically incredible, think about the psychological weight of a medical record that never sleeps. If your health data is never truly finalized, and the answers to your biological mysteries could arrive tomorrow, or in 5 years or in a decade, what happens to the concept of closure?
19:24In a world of proactive genomic reanalysis? We might have to accept that our medical stories never actually end. They are just perpetually waiting for the next software update. Will this continuous reanalysis become a fundamental human right in standard healthcare, or will it just be a premium service for those who can afford it?
19:40That's the real question. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating.
19:56If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
20:05Thanks for listening and join us next time as we explore more science, base by base.