This episode reviews a systematic analysis of ultra-rare FBN1 variants in the 100,000 Genomes Project using SpliceAI, RNA assays and minigene tests. The study identified 20 non-canonical splice variants across 23 families, confirmed splicing defects for 16 variants, and estimates these variants account for ~3% of undiagnosed FTAAD/Marfan families. The work highlights the value of intronic analysis and confirmatory RNA testing in clinical genomics.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So I want you to just imagine the human aorta for a moment.
0:12Oh, yeah, the main highway of the body. Right, exactly. It's this massive vital vessel, and it's carrying all that oxygen much blood directly from your heart out to the rest of your body. And I mean, it has to withstand incredible hydrostatic pressure.
0:25Oh, absolutely. It's absorbing the shock of every single heartbeat every 2nd of your life. Yeah. So now, consider what happens when a tiny spelling mistake, you know, something hidden deep within the so-called junk DNA of a patient's genome threatens the structural integrity of that exact vessel.
0:43It's I mean, it basically creates a ticking time bomb inside their chest. It's a really terrifying clinical reality. It really is. And we are looking at families who suffer from these life-threatening aortic aneurysms that they stretch across multiple generations.
0:58Right. Their aort is just progressively weaken and dilate under all that pressure. So they're carrying this constant looming risk of a sudden tear or a dissection. And the really agonizing part of this diagnostic odyssey is that these families, they know something is terribly wrong, right?
1:18Like they see the family history. Oh, for sure. They go to the clinic, they do the standard genetic test, hoping to find the cause so they can, you know, figure out a clinical strategy. But then the result come back completely blank.
1:29They are literally told their genetic tests are normal. Yeah, and the tragic reason those results come back normal isn't because the mutation isn't there. It's simply because the tests are, well, they're looking in the wrong place entirely.
1:40Exactly. So how could this change the fate of families who have been told their genetic test is normal when their medical history screams otherwise? Well, the standard sequencing pipelines? They focus almost exclusively on the coding regions of our DNA, right?
1:55The Exxons. They largely ignore the vast dark spaces of our genome. Introns. Right? The intervening non-coding regions. But we were finally reaching a point where artificial intelligence is turning the lights on in those dark spaces.
2:08Which is incredible. And before we get too deep into how the AI does that. Today we celebrate the work of the researchers at Genomics, England Limited, the universities of Oxford and Manchester, and the incredible 100,000 Genomes Project team, who have advanced our understanding of MarFans syndrome and connective tissue disorders.
2:26Yeah, to really grasp the magnitude of what they did in this deep dive. We have to talk about Marfent syndrome. It was 1st described way back in 1896, I believe. Yeah, 1896. And it affects roughly 2 to 3 in every 10,000 people globally, which, when you hear Marfan syndrome, I feel like your mind usually goes straight to the physical stuff.
2:46Right, like the extreme height or the disproportionately long limbs. Exactly. Skeletal anomalies, maybe Chen, we abnormalities or, you know, lens dislocation in the eyes as outward phenotypes. But while those signs are prominent, the true threat.
3:00The thing that demands urgent clinical attention is cardiovascular. Specifically, familial thoracic aortic aneurysm disease or FT. FTA. Right. And without proactive surgery, this just progresses relentlessly, doesn't it?
3:16It does. The connective tissues simply cannot hold up to the mechanical stress. And we've known the primary genetic culprit for a while now. It's a gene called FBN1. FBN one, and that encodes a protein called Fibralin one, right?
3:29Exactly. It acts like molecular scaffolding in your connective tissues. And when standard tests sequence the FBN1 gene in patients with a clear MarFan diagnosis. They find the variant in over 90% of cases.
3:41Over 90%. So let's unpack this for a second. Standard genetic testing is essentially like proofreading a book by only looking at the actual text of the chapter. That's a great way to put it. But what if the critical typo is written in the blank margins of the pages?
3:54Because current clinical tests, they typically stop proofreading just like 8 letters past the chapter boundary. The Exxon boundary. Yeah. So they completely miss these cryptic, non-canonical splice variants hiding deep in the margins.
4:09But here's my question. If 90% of cases are solved easily. Why go through the massive trouble of digging through all that entronic junk for the remaining fraction? Is it really worth it? Oh, it's absolutely worth it.
4:22It's a life or death value for those specific families, without a confirmed genetic variant on file. Clinical teams can't perform cascade testing. Cascade testing being when you screen the asymptomatic relatives.
4:34Exactly. Kids, siblings, cousins. You need to see who inherited the Riskalil, and who didn't. without a known variant, you can't test them. You're just waiting around for a structural failure to show up on an ultrasound or an MRI.
4:46Right, you're leaving the whole extended family in this terrifying clinical limbo. Finding that hidden typo fundamentally alters the trajectory of an entire family tree. It triggers life-saving cardiac surveillance for the carriers and releases the non-carriers from a lifetime of anxiety.
5:00Okay, that makes total sense. But to find a type of that rare, you need a massive haystack to search through, which brings us to the data set they used. And it is a massive data set. They tapped into the 100,000 genomes project.
5:13Yeah, they started with aggregated data from 78,195 individuals, and then they filtered that down to a really specific group of 703 people. Right, 703 individuals who were recruited specifically because they had FT.
5:28That filtering is critical. These people had already gone through standard testing and came up empty. They were the unsolved mysteries. Exactly. And to find those hidden variants deep in the introns. They needed a completely new kind of tool.
5:40So they turned 2 splice AI. Splice AI. Now, I know this is a deep neural network used to predict splicing errors, but how does an AI spot a mistake that a human geneticist, you know, using standard algorithms would just completely miss?
5:53Well, traditional algorithms rely on rigid, human coded rules. They look for specific sequences right at the boundary between the Exxon and the Intron. But Spice AI doesn't do that. It was trained on massive data sets of RNA transcripts to learn the really subtle, long-range context clues.
6:12The context clues that the actual cellular machinery uses. Yes, the split system. The AI recognizes the broader genomic landscape, not just the immediate boundary. And what's crazy is the researchers made a huge adjustment here.
6:26Usually the search window for this kind of AI is just 50 base pairs, but they expanded the AI's view to 500 base pairs. Why do that? Because when you're hunting for deep intronic variants, you're looking for sequence anomalies that set 100s of bases away from the normal cut sites, a 50 base bare window is too narrow.
6:43By widening it to 500, the AI can evaluate long-range disruptions that a narrower search is entirely blind to. Wow, okay, so it casts a much wider net. But here's where it gets really interesting to me.
6:54And AI prediction is still just code on a screen, right? You can't base a major heart surgery on a computer prediction. No, absolutely not. You need physical proof. You have to prove the patient cells are actually misreading the RNA in the exact way the AI predicted.
7:09And they used RNA testing from blood samples to do this. RTPCR and RNA sequencing. But wait, FBN1 is a connective tissue gene. It builds structural scaffolding. Blood is not connective tissue. It barely expresses in the blood.
7:23So how on earth do you extract meaningful RNA data from a simple blood draw to prove an AI right? It is incredibly difficult. FBN one expression in whole blood is just a whisper against a deafening background of other genes, extracting that low abundance transcript without missing the variant requires serious laboratory expertise.
7:43It's like finding a needle in a haystack inside another haystack. Exactly. The clinical lab team had to design highly specific primer sets just to capture that faint signal reliably. The fact that they managed to validate these from standard blood draws is a huge operational victory.
7:57Because the alternative is what? Taking a skin biopsy? Yeah, taking a skin punch biopsy and spending weeks culturing cells. Blood is so much faster and less invasive. And for patients who were deceased or couldn't provide fresh samples, they actually used minagene constructs.
8:13Many genes. I love this part. They are basically lab made miniature versions of the gene, right? Yeah, they custom build a miniature version containing the Entronic sequence and the patient's unique variant.
8:24Then they put it into cultured lab cells to see if the cells splice it incorrectly. It's like a biological sandbox to test the AI. So brilliant. Okay, let's talk about the numbers because the findings are staggering.
8:35Out of almost 14,000 variants in the FBN one gene across all those people, the AI flag just 21 is dangerous. Just 21. And the enrichment of those variants is incredible. They calculated an odds ratio of 84.
8:50Wait, 84. Just to clarify for everyone listening. What does an odds ratio of 84 actually mean? It's an astronomically high signal. An odds ratio of one means the variant is equally likely to be in a healthy person as a sick person.
9:0384 means it's 84 times more likely to be found in the FTA group than the general cohort. Wow. So the AI really cut through the noise. Completely. In total, they found 20 unique cryptic variants across 32 individuals in 23 different families, and a staggering 70% of these variants lay far beyond the standard testing boundaries.
9:2570%. So they were way out in the deep margins. And the mechanism here is fascinating. How does one altered letter out in the junk DNA actually destroy a protein? The study mentioned that 9 of these variants created a pseudoexon?
9:40Right, pseudoexxons. So normally the splices on the sales machinery, scans the RNA, finds the start and end of an intron, and cucks the whole junk section out. Stitching the real Xxons together. Exactly.
9:50But these deep intronic mutations trick the machinery. They all do the sequence just enough to look like a real Exxon boundary. Like a mirage in the desert. Yes. The cell gets tricked into reading junk DNA as a real Exxon, and it accidentally leaves a chunk of intronic DNA inside the final transcript.
10:08That's the pseudo-exon. And when the cell tries to build the fibrillin one protein from that transcript? It hits a stretch of absolute nonsense. It usually causes the cell to just degrade the RNA entirely or it makes a broken protein that ruins the connective tissue.
10:22Man. And you can see the real world impact of this. Look at family 22 from the paper. This is a 4 generation family with a history of aneurysms. And finding the single variant changes everything for them.
10:34It really does. Because of this discovery, 10 to 20 relatives in family 22 can now get targeted cascade testing. Meaning they finally know exactly who needs the preventative cardiac care. Exactly. It's lifesaving.
10:46But you know, the mechanisms aren't always just inserting a chunk of nonsense. Family 9 was a really fascinating anomaly. Oh, right. Family nine. Their variant wasn't deep in an intron. It was at the very beginning in the 5 Prime UTR.
10:58Yeah, the 5 prime untranslated region. It's basically the landing pad for the cellular machinery before it starts building the protein. So what did the mutation actually do there? It messed up the splicing and created an upstream open reading frame or a URF.
11:12A URRF. So it acts like a false start signal. Yes. When the machinery lands, it hits this false start line first, it starts building a short, completely irrelevant protein, which totally disrupts its ability to reach the actual FBN one gene just downstream.
11:28So it's like creating a massive traffic jam right on the on ramp, so the cell can never actually get onto the highway to build the real protein. That is a perfect analogy. The cell is starved of fibrill and one.
11:39So what does this all mean for the families? They finally get a diagnosis, but did the AI predictions actually hold up in the lab for most of these? They absolutely did. RNA testing physically prove the average splicing ends 16 out of the 20 variants.
11:5416 out of 20. That's a massive success rate for an AI prediction deep in the junk DNA. It really is. Yeah. And when you look at the implications, this deep intronic analysis increases the diagnostic yield by about 3% in unsolved FTAD cases.
12:093% might not sound like a huge number to some people, but for the families on that diagnostic odyssey, that 3% is everything. It's the end of years of anxiety. Exactly. And because of this, the researchers are actually suggesting that we change the clinical rule book.
12:23Specifically, the ACMG guidelines. The American College of Medical Genetics. Right now, when labs find a deep entronic variant, it usually just gets labeled a VUS, right? A variant of uncertain significance.
12:35Yeah, genetic purgatory. Doctors can't really act on a VOS, but the researchers propose elevating these deep intronic variants to PBS one strong. PVS one strong. What does that mean, practically speaking?
12:48It means pathogenic. Very strong. If CDNA analysis physically proves the variant disrupts crucial functional domains of the protein, elevating it to PBS one strong gives doctors the green light to officially diagnose the patient and start clinical interventions.
13:02That's huge. It cuts through the red tape. But I do want to touch on the limitations of the study, because science is never perfectly neat. They tried to replicate these findings in the UK biobank, right?
13:12They did. They looked at the UK biobank, which is massive. But they didn't find the same enrichment for generic aortic aneurysm variants. Why not? Did the AI miss something? No, it's about how the data is categorized.
13:24The 100,000 genomes project had highly curated patients with the rare genetic FTA. But the UK biobank uses general hospital billing codes. Ah, so their code for aortic aneurysm mixes genetic cases with like aneurysms caused by age or lifestyle factors.
13:42Exactly. Smoking, hypertension age. The underlying biology is completely different. It just proves how specific this genetic mechanism really is. You need precise clinical data for the AI to work effectively.
13:53That makes total sense. So let me ask you this. If RNA testing is the gold standard to prove these AI predictions. Why isn't every clinic doing it right now? Are we hitting a bottleneck? We are definitely hitting a bottleneck.
14:03It's a heavy resource burden. Like we talked about, getting fresh patient blood samples quickly before the RNA degrades is tough. Logistically, yeah, that's a nightmare for an average clinic. And clinical labs really need to scale up their splicing analysis capabilities.
14:17The AI can find the target in seconds. But the physical lab validation still takes immense skill and time. Right. So just to pull all this together for you listening by pairing deep neural networks like Splice AI with advanced RNA laboratory testing, scientists can finally illuminate the dark, non-coding regions of the genome.
14:37Yeah. This approach uncovers cryptic splice variants that standard tests completely miss. Exactly. And it directly ends diagnostic odysseys for around 3% of unsolved MarFan syndrome cases, which triggers life-saving cardiac care for whole families.
14:52This is total game changer. It really is. And it leaves me thinking, what does this mean for all the other rare, unsolved genetic diseases out there? I mean, if expanding our view by just a few 100 base pairs into the genomic junk can save families from setting heart failure, what else is hiding in the dark spaces of our DNA, just waiting for the right AI to find it?
15:13That's the 1000000 question. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
15:29If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
15:38Thanks for listening and join us next time as we explore more science, base by base.