Genome sequencing of 78,195 participants in the 100,000 Genomes Project identified ultra-rare non-canonical FBN1 splice variants enriched among individuals recruited with familial thoracic aortic aneurysm disease (FTAAD). Experimental RNA assays confirmed aberrant splicing for most candidates, including multiple deep intronic pseudoexon events, indicating a measurable diagnostic contribution from intronic variants beyond standard clinical testing windows.
0:15Welcome to Base by Base. We're here to really get into the details of some fascinating science. That's the plan. Taking complex research and breaking it down. And today, we're looking at MarFan Syndrome.
0:25Now, you might know it as a genetic condition affecting connective tissue, but finding the precise genetic cause, well, that's not always simple. No, it really isn't. Even when someone clearly has the clinical signs, the standard genetic tests sometimes come back.
0:40Well empty. The answer can be hidden. Hidden somewhere, standard tests, just don't look. And that's exactly what a recent paper tackles. It's research led by Susan Walker and her colleagues published in genetics and medicine just this year, 2025.
0:54Right. And their work really dives deep into Marfan, hunting for those genetic changes that have been missed. Our goal here is to walk through their study, understand their methods, the results, what it all means.
1:06We want to see how using things like whole genome sequencing, reading the entire DNA instruction book, not just the main chapters, plus some smart computer analysis and crucially, lab work can uncover these hidden genetic causes.
1:20specifically variants that mess up something called splicing. Exactly. Splicing is a key process we'll definitely get into. Okay, let's start with the basics. More fan syndrome, MFS. It's a connective tissue disorder, affects maybe 2 or 3 people in every 10,000.
1:35And while it can be diagnosed at different ages, it often becomes apparent in the late teens or early 20s, connective tissue is, you know, everywhere in the body is the glue, the framework. So if that's faulty, it can cause problems all over.
1:48People might be tall, have long limbs, maybe eye issues. Yes, those are common features. But the really serious concern, the life-threatening part is often the big artery coming from the heart. You can weaken and stretch, leading to what's called familial, thoracic, aortic, aneurysm disease or FTT.
2:04And that can get to a tear, a dissection. Precisely, which is incredibly dangerous. That aortic risk is why getting an accurate diagnosis is so vital. Doctors use clinical criteria, the Ghentanosology to assess the signs.
2:18Now, genetically, we've known for a while that the main player is a gene called FBN1. Right, FBN one. It's a big gene, huge, actually. Over 237,000 DNA letters broken into 66 exxons, which are the bits that code for the protein.
2:34And that protein is fibrill and one. Part of that connective tissue scaffolding we mentioned. Exactly. Fibrillin one is crucial for the extra cellular matrix. So if the FBN one gene has a typo, the instructions for building that scaffolding are wrong.
2:47Make sense. And we know lots of different types of typos or variants in FPN1 can cause MarFan things that stop the protein being made too early, small insertions or deletions, single letter changes. Especially changes affecting specific building blocks called cystines, which are really important for the protein shape.
3:04And interestingly, different changes in FBN one can sometimes cause related but distinct conditions too, not just classic MarFan, like galeophysic dysplasia. So it's complex. But here's the puzzle. Even with good FBN one testing, over 90% of people with clear clinical MarFan get a genetic diagnosis, but there's still that remaining group, maybe 5, 10%, where the tests are negative.
3:28Yeah, and these families are often on with the paper calls a long diagnostic odyssey. They know something is wrong, often affecting multiple family members. But the standard genetic tests don't provide an answer.
3:39Why not? are the possibilities? Well, maybe the clinical picture isn't quite MarFan. Maybe it's another gene altogether. But a big reason, and the focus here is these cryptic variants within the FPN one gene itself, variants that hide from standard tests.
3:54Cryptic variants. We hear that term. Sometimes it means big rearrangements, but this paper is focused on something else, right? Cryptic splice variants. Exactly. These are often small changes like single letter swaps, but the key is where they are.
4:05They're not in the Exxons, the coding bits. And crucially, they're not of the tiny bit of intron. The non-coding stuff right next to the Exxons that tests usually check. Maybe the 1st 8 letters or so. They're buried deeper in those long intron stretches.
4:17Precisely. Deep intronic variants. Or sometimes in other regulatory regions, and standard tests, which focus on Exxons and those miniate flanking bases, just fly right over them. They don't see them. And what do these variants do?
4:31You mentioned splicing. Right. So think of splicing as the cells film editor. The gene gets copied into a long RNA molecule, including both the exxons, the scenes you want, and the introns, the bits you don't. Splicing cuts out the introns and stitches the exxons together perfectly to make the final moody script, the MRNA that tells the cell how to build the protein.
4:52Okay, the editing process. Yeah. And a cryptic splice variant acts like a bad edit instruction. It might create a new, wrong place to cut or join or hide the correct place. It messes up the editing, leading to a garbled MRNA script, and usually a protein that doesn't work properly.
5:07So Walker and colleagues set out to systematically figure out how common this problem is, these hidden splice variants causing MarFan or FT in those undiagnosed cases. They leveraged a fantastic resource.
5:18The UK's 100,000 Genomes project, the 100 KGP. Ah, yes, that massive project using whole genome sequencing for rare diseases. Exactly. It was designed specifically for situations like this. Families with suspected genetic conditions, but no answer from standard test.
5:35They looked at genome data from huge number of people in the project. How many people? Their initial analysis used aggregated data from over 78,000 individuals, and critically, within that group, there were over 700 people recruited because they had FTAD, that thoracic aortic disease often linked to Marfan.
5:53FTAD was actually one of the more common reasons people joined the rare disease part of the 100 KLGP. So they had a large group with the key clinical feature and a very large comparison group. What was the strategy?
6:04They systematically scan the FBN one gene in everyone, looking for ultra rare variants. Initially, they focused only on singleton's variants seen just once in that entire data set of nearly 80,000 people.
6:16The idea being that such rare changes are more likely to cause a rare disease. And how do they know which of those rare variants might affect splicing? They used a tool called Splice AI? Yes. Splice AI was key.
6:27It's a computational tool, a type of AI, really. It's been trained on vast amounts of genetic data to predict how likely any given DNA changes to mess up splicing. How does it work? What does it tell you?
6:39You feed it the variant, and it looks at the DNA sequence around it. Then it gives you scores, from 0 to one. A high score suggests a high probability that the variant will either create a new splice site where there shouldn't be one game or break an existing normal one, loss.
6:55It predicts different types of splice site changes. Okay, so find the ultra rare variants, run them through splice AI, and then C. Do the ones predicted to affect splicing show up more often in the FTAid group?
7:06That's exactly the question. Is there an enrichment? If those predicted splice variants are significantly more common in people with FD, that's strong evidence they're actually contributing to the disease.
7:15So what did they find in that 1st big scan? Okay, get this. Across those 78,000 genomes, they found almost 14,000 different ultra rare singleton variants just within the FBN one gene sequence. 14,000 in one gene.
7:30Wow, that's a lot of rare variation. It really highlights how much variation exists between individuals, but here's where splice AI comes in. Out of those nearly 14,000 variants, only 21 had a splice AI score, suggesting a potentially moderate to strong impact on splicing.
7:45They used a cutoff of .5 or higher initially. Just 21 candidates flagged out of 14,000. That really narrows it down. where did those 21 turn up? This was the crucial finding. They looked at the frequency.
7:56In the 703 individuals recruited with FDAid. 9 of them carried one of these 21 potential splice variants. That's about one.3% of the FT group. Okay. And in the others, the big comparison group. In the over 77,000 people without an F tape recruitment diagnosis, only 12 individuals had one of these variants.
8:13That's an incidence of just .015%. Right. So one. 3% versus .015%. That sounds quite different. It's hugely different. Statistically, the odds ratio was 84. That means someone in the FDA group was 84 times more likely to carry one of these predicted splice variants than someone in the comparison group.
8:30And the P value was incredibly small. 9.7 times 10 to the -14, meaning this association was extremely unlikely to be due to random chance. An odds ratio of 84. Yeah, that's a really powerful signal telling you these variants are strongly linked to the disease.
8:45Absolutely. strong statistical evidence. Did they look at different splic AI cutoffs? Because .5 is just one choice. They did. And that's important. As you change the threshold, you change the balance.
8:56A higher threshold, like 0.8 is stricter. You catch fewer variants, only 10 here, but a higher proportion of them were in the FDAD cases, 7 out of 10. A lower threshold, like .3 is more sensitive. It catches more potential variants, 34, including 11 in the FTED group.
9:14The enrichment was significant across different thresholds. But exploring this helps understand the tradeoffs between finding more potential hits versus getting fewer false positives. And did they also look at variants that weren't quite so rare, not just singletons?
9:28Briefly, yes. They relaxed the criteria slightly to include variants seen up to 5 times. They still found significant enrichment in the FTAD group, but the odds ratio was lower, around 35, which makes sense.
9:40Slightly more common variants are less likely to be the sole cause of a rare disease. Okay, so that initial scan gave really strong statistical backing, but they wanted to find more families, right? To maximize the diagnostic potential.
9:52Exactly. So they did a secondary analysis, focusing just on the full set of families, recruited into the 100 keyDP rare disease programs specifically for F Day. This let them include families missed in the 1st aggregated data set, maybe due to data access timing, or using older genome references.
10:09This focused cohort included 672 families with F Day. And they tweaked their search strategy for this deeper dive. Yeah, they broadened the net a bit. They didn't restrict it to only singleton variants anymore.
10:22They allowed variants with a frequency up to 0.one% in the population. They also use a slightly lower splice AI threshold. A .2 would be more sensitive. And they pulled in data from families analyzed earlier, plus one family identified through routine clinical testing using similar software.
10:37What did this more targeted, wider search turn up? It was successful. This secondary analysis identified another 14 families who had candidate FBN one splice variants meeting these criteria. There were 11 different variants found across these 14 families.
10:51So put it all together, the initial scan, this secondary analysis. What's the final count? In total, combining all phases and including a few known variants from previous reports in these cohorts, they identified 20 unique candidate splice variants.
11:06These were found in 32 individuals across 23 different families, all previously lacking a genetic diagnosis, despite having features of MarFan or FTED. 23 families getting a potential genetic answer. They didn't have before.
11:20That's really impactful. Where did these 20 variants actually occur in the FPN1 gene? They were pretty spread out, located in introns ranging from intron one right near the start of the gene, all the way out to intron 63 near the end.
11:32There wasn't one single major hotspot, although a couple of introns did have variants found in more than one family. And the key question, were they mostly in those cryptic regions? Yes. This really validated the whole approach.
11:44Of the 20 you need variants identified 14 of them. So 70% were located outside the standard plus minus 8 base pair region around the Exxons that typical tests analyze. They were indeed deep intronic variants, confirming they would have been missed by conventional sequencing.
11:59Okay. And what did Splice AI predict these 20 variants would actually do to the splicing process? There were a few different predicted outcomes. The most common scene for 9 of the 20 variants was something called pseudo-exanization.
12:13Pseudoexonization. creating a fake Exxon, like adding a scene that doesn't belong. That's a great analogy. The variant creates splicing signals within an intron that trick the cell's machinery into including a chunk of that intron sequence in the final MRNA message, as if it were a real Exxon.
12:31This inserts junk information, basically, and almost always messes up the resulting protein. What else did Splice AI predict? Exon Extension was also common, predicted for another 9 variants. This is where the normal splice site is ignored, and a cryptic site further into the entron is used instead, making the adjacent exxon longer than it should be.
12:49One of these even affected the non-coding 5 Prime UTR region at the start. And less common predictions. They saw one predicted case of Exxon skipping where a whole real Exxon is just left out, and one case of Exxon contraction, where only part of a real Exxon is included.
13:05Interestingly, none of the variants they found in this set were standard coding variants that also happened to affect splicing. These are primarily variants disrupting splicing from within the non-coding regions.
13:18Now, predicting this with a computer is one thing, but you really need to prove it happens in the cell, right? Experimental validation. Absolute critical. You can't just rely on the prediction, especially for clinical diagnoses.
13:29You need to show that the variant actually causes abnormal splicing in patient cells or a model system. How do they try to validate these predictions. They looked at RNA sequencing data. They did. RNASEC data was available for 10 of the 23 families.
13:43RNASIC basically reads out all the RNA messages in a cell sample. And did it work? Did it confirm the splicing errors? Mostly. No. It turned out to be quite challenging for FBN one. The problem is that FPN one is expressed at very low levels in blood cells, which was the sample type available for most people in the 100 KGGP.
14:00Because there wasn't much FBN one RNA to begin with, the RNA sec didn't capture enough information. Across all 10 data sets, they found only a single RNA red that supported one of the predicted splicing events.
14:13Wow, just one red. So RNASUC wasn't the magic bullet here for validation. Not from blood samples, no. It really highlights that the best validation method depends on the gene and the available tissue. Their main workhorse for validation was actually RTPCR.
14:28Right. Right. Reverse transcript taste PCR. Explain that again. You take RNA from the patient's blood, convert it into a more stable DNA copy, ZDNA, and then use PCR primers designed specifically to amplify the region around the predicted splice defect.
14:41If abnormal splicing is happening, you'll see PCR products of unexpected sizes, which you can then sequence to confirm exactly what went wrong. And was RTPCR more successful? Yes, much more so. Despite the low expression challenge, they managed to get informative results, using RTPCR from blood RNA, for every case they attempted, 11 families used it as primary validation, +2 others for confirmation.
15:06It sometimes needed careful optimization, but it worked, and was less resource intensive than other methods. The main limitation was needing a frag blood sample. They also use something called minogene essays.
15:17They did for 3 families. Menagene essays where you build a small artificial gene construct in the lab. It contains the Exon and surrounding intron sequence from FBN one, both the normal version and the version with the patient's variant.
15:30You put these constructs into lab grown cells, which then transcribe and splice the minagene. And you compare the results from the normal versus the variant construct. Exactly. If the varying construct produces abnormally spliced RNA compared to the normal one, it confirms the variance effect on splicing.
15:45They use this for 3 families and it confirmed the predictions in all three. It's useful if you can't get patient RNA, but it takes more time and resources than RTPCR. And designing constructs can be tricky sometimes.
15:56So adding it all up. How many of the 20 candidate variants got experimental confirmation? They achieved experimental validation for 16 out of the 20 unique variants. 15 were confirmed in this study through RTPCR or menagenes, and one recurrent variant had been previously confirmed and published by another group.
16:1616 out of 20 validated. That's a pretty high success rate. It really backs up the idea that their computational approach was flagging real culprits. It does. It shows that combining genome sequencing with careful en silico prediction and targeted experimental follow-up is a powerful strategy for finding these cryptic variants.
16:36Okay, so they found them and validated many. But the paper goes into quite a bit of detail about interpreting the splice AI results, suggesting it's not always straightforward. It's more than just looking at that main score.
16:47Right. This is really important for getting it right clinically. Just taking the default splice AI delta score, which measures the change in predicted splicing caused by the variant at face value, isn't always enough for accurate interpretation or even for designing the validation experiments.
17:01What kind of nuances did they find? They mentioned window size. Yes. Spice AI, by default, looks for potential spice sites within a relatively small window around the variant, usually 50 base pairs. But these cryptic variants can sometimes activate or affect splice sites that are much further away.
17:18The researchers found that expanding the analysis window, say to 500 base pairs, sometimes gave a more accurate prediction. Can you give an example? Sure. Family 17 had a variant, see 1837 +5 GC. The standard 50 BP window predicted only a moderate effect.
17:35But when they looked further out in a 500 BP window, Splice AI gave a very strong prediction for a new Splice site being created about 60 base pairs away. And did the lab work confirm that? It did. RTPCR showed that the Exxon was extended by 66 base pairs consistent with that more distant site being activated.
17:54So if they'd only looked to the default window prediction, they might have underestimated the variance impact or struggled to design the right PCR primers to see the effect, looking wider was key. They also emphasize looking at absolute splice AI scores, not just the change, the delta score.
18:08Why is that important? This is a bit more subtle. The Delta score tells you how much the variant changes the score. But the absolute score tells you the inherent likelihood of a particular sequence being used as a splice site, whether it's the normal site or a potential cryptic one.
18:24Sometimes a cryptic site might already exist in the normal DNA and have a reasonably high absolute score, even if it's not usually used. Okay. Now, a variant might only increase that cryptic site score by a small amount giving a low delta score.
18:36But if that small increase pushes the absolute score of the cryptic site higher than the absolute score of the nearby normal spice site, the sales machinery might suddenly prefer the cryptic site. Ah, so the variant just tips the balance even if the change itself wasn't huge.
18:50Exactly. The example from family 23, C.4942 plus 1980, show this well. Low delta score, .2. But the variant pushed the absolute score of a nearby cryptic donor site from .79 up to .99, making it stronger than the normal site, which drops slightly to .84.
19:08And the functional result. RTPCR confirmed. It caused a 13 base pair extension of the Exxon. This variant was previously considered likely benign, probably based on just the low Delta score. Looking at the absolute scores revealed its true potential.
19:23It even helped them figure out what was going on in Family 7, where initial experiments failed until they re-examined the absolute scores to design better primers. So these deeper computational analyses are vital for correct interpretation and successful validation.
19:36It's not just plug and play. Not at all. It requires careful examination to avoid misclassifying variants or missing effects. It really bridges the gap between finding something suspicious and understanding if it's truly pathogenic.
19:51Let's shift to the consequences for the protein. You said 13 of the 20 variants were predicted to cause a loss of function, probably by creating premature stop codons. Yes, that's the most common way FBN one variants cause more fan.
20:03It leads to Apple insufficiency, basically. Having only one working copy of the gene isn't enough to make the required amount of fibrile in one protein for healthy tissues. The faulty message often gets destroyed, or a useless shortened protein is made.
20:16But what about the other 7 variants? The ones that didn't seem to create stop codons? These are often trickier. They resulted in splicing changes predicted to keep the protein sequence in frame. That means they might insert extra amino acids from a pseudoexon, delete some amino acids from exxon skipping or contraction, or affect non-coding regions. For these, just knowing the splice defect happened isn't always enough.
20:40You need to think about what those changes do to the protein structure or function. So adding or removing a few amino acids might matter a lot or maybe not. Exactly. Depends where in the protein it happens.
20:51Is it in a critical structural part? Does it disrupt how the protein folds or interacts with other molecules? Some variants, like one causing a tiny 2 amino acid insertion, family 6, are really hard to predict without functional studies.
21:04Others involve larger in frame insertions? families 11 and 19, or deletions, family 17, 18, which are more likely to be disruptive but still need careful assessment. And there was that unusual one in the 5 prime UTR, the non-coding start region.
21:17Yeah, Family 9s variant, C 182 +one GA. It affected the very 1st non-coding Exxon, extending it by just 4 DNA letters, but those 4 letters created a new AUG start code on before the real start code on for the FBN one protein.
21:33Creating an upstream open reading frame, a UORF, you said that can act like a decoy. That's right. The cell's protein making machinery might start translating at this fake start site, which can then interfere with it efficiently finding and using the real start site downstream.
21:46The hypothesis is that this reduces the overall amount of normal fibrile in one protein being made, even if the main coding sequence is okay. That's a subtle way to cause disease, regulatory almost. It is.
21:58And confirming that mechanism would need specific lab assays. But the family members had more fan features and negative standard testing, making it a compelling candidate. These regulatory variants are definitely underappreciated.
22:09Did any variant appear in multiple families? Yes, one really stood out. C1589 through World 217 GT. They found this exact variant in 7 different people from 4 unrelated families in the UK. Wow, a recurrent cryptic variant.
22:24Was it Splice AI score high? Interestingly, no. Its initial score was relatively low. 0.20. But the fact that it kept popping up independently in unrelated families with Marfin F Todd was a huge red flag, suggesting it was truly pathogenic.
22:39Did they check if the families might be distantly related, sharing an ancestor? They did hepotype analysis, looking at the genetic background around the variant. The results suggested these families were not related recently.
22:50It seems this specific spot in the intron is just prone to this particular mutation, occurring multiple times, independently, a mutational hotspot. So this highlights that even variants with lower prediction scores need attention if they recur in affected individuals.
23:06Absolutely. Its recurrence, combined with a previous publication showing it creates a 202 base pair, pseudoexon, makes it a clear pathogenic variant that labs should probably be specifically screening for it, even though it's deep and chronic.
23:19Thinking about the overall impact for the patients, what percentage of these previously unsolved F-Tad cases got an answer through this approach. What was the diagnostic yield? Within this specific cohort of families in the 100 TGP, who had FT features, but no prior genetic diagnosis.
23:34Finding these non-canonical splice variants solve the case for about 2.8% to 3.3% of them, depending on the analysis set. Around 3%. That might not sound massive, but for those individual families after years on that diagnostic odyssey.
23:48It's everything. It really is. These were families who'd often had multiple rounds of testing gene panels, eggzomes that came back negative. Finding the specific FPN one splice variant was the breakthrough.
24:00And clinically, did these families have classic Marfan features? It was a spectrum, which is typical for MarFan. They had various combinations of the skeletal signs, eye problems like lens dislocation, and critically the aortic involvement.
24:12Aortic roof size varied quite a bit from near normal and some younger relatives to requiring major surgery in others. And the power of getting that specific genetic diagnosis is cascade testing, right?
24:22Yes, that's a major benefit. Once you know the exact variant causing the condition in one person, you can efficiently test their relatives, parents, siblings, children. Those who inherited the variant can then get proactive cardiac surveillance, monitoring their aorta, maybe starting medications, planning interventions if needed.
24:40Those who didn't inherit it can be reassured. It directly guides potentially life-saving management. The paper even shows pedigrees illustrating this impact. How did these variants end up being classified clinically?
24:52Pathogenic, uncertain. Using the standard ACMGMP guidelines, adapted by the Klingon expert panel for splicing, they classified 9 of the 20 unique variants as pathogenic, 5 is likely pathogenic, and 6 remained as variants of uncertain significance, PUS.
25:08Of course, VOS classifications can sometimes be upgraded later if more evidence emerges. Now, the study also tried to replicate this in the huge UK biobank data set. How did that go? Right. UK Biobank has genetic data on almost half a 1000000 people, linked to health records via diagnostic code like ICD 10.
25:25They looked for participants coded with aortic aneurysm dissection, I 71, or specifically Marfon Syndrome, Q87.4. Did they see the same enrichment of these predicted FBN one splice variants in the aortic group by 71.
25:39as they saw in the 100 kilo GP FTAD cohort. Interestingly, no. In the Braun I 71 group in UK Biobank, there was no significant enrichment. They found some rare predicted splice variants, but they weren't statistically more common in people with that code compared to the rest of the biobank population.
25:56Why the difference? The likely reason is phenotype specificity. The I 71 code is very broad. It includes abdominal aortic aneurysms, which are much more common than thoracic ones, and usually related to things like smoking and high blood pressure, not typically single FBN one variants.
26:10The 100 keto GP cohort was specifically enriched for familial thoracic disease, a much better fit for FBN one involvement. The UKBI 71 group was just too diluted with other causes. But what about the UKB participants coded specifically with MarFans syndrome, Q 87.4.
26:27There they did see a very strong enrichment. The predicted splice variants were significantly more common in that group. with a huge oz ratio, similar to the 100 D to GP findings. So that confirms it then?
26:38It supports it. But with a significant caveat the authors mention, it's quite possible that many people got the Q 87.4 code, because they'd already had clinical genetic testing that found an FBN one variant, maybe even one of these splice variants.
26:52So finding the variant, again, in the UKB data, for someone already diagnosed based on that variant, isn't truly independent confirmation, it could be circular reasoning. I see. So the UKB results support the link and clinically diagnosed Marfan, but the broad aortic code wasn't specific enough.
27:08It really highlights how crucial detailed clinical information is. Absolutely. You need good phenotyping, even with massive data sets. Okay, let's pull back and synthesize. What are the big takeaways from this study by Walker and colleagues?
27:21The clearest message is that these non-canonical splice variants, the ones hiding deep in the introns are a real and probably underdiagnosed cause of our fan and related aortic disease. They're not just theoretical.
27:33They explain about 3% of previously unsolved cases in this well-defined co-work. And finding them requires whole genome sequencing. Yes, that's the enabling technology because it reads the introns. Standard Exon sequencing our panels will miss them.
27:46But the sequence alone isn't enough. You need the whole package. Exactly. You need the genome sequence data. You need sophisticated prediction tools like Splice AI, used intelligently considering things like analysis windows and absolute scores.
28:00And then, crucially, you need experimental RNA validation to prove the predicted splicing defect actually occurs in the patient's cells. It a multi-step process. What does this mean for clinical genetic testing for MarFan going forward?
28:14The author strongly suggests that testing should incorporate analysis of introns looking beyond the standard splice boundaries. And when candidate splice variants are found, labs need pathways for confirmatory RNA testing.
28:26What are the roadblocks to making that standard practice everywhere? There are definitely challenges. Clinically validated pipelines for analyzing introns systematically aren't universal yet, and setting up reliable clinical grade RNA testing, especially for low expression genes like FBN one from blood, require specific expertise, resources, and sample handling protocols that many labs might not have readily available.
28:51It's a significant logistical hurdle. How might this influence the official guidelines for classifying these variants? The paper touches on the Klingon FPN one guidelines. While they cover standard splice variants well, They could be clearer or perhaps stronger regarding evidence from experimentally confirmed deep intronic splicing events, like pseudo-exon inclusion or complex in frame changes.
29:11Strengthening criteria like PVS one for demonstrated aberrant splicing could help resolve more of US classifications. Which would provide more clarity for patients and doctors. Exactly. Reducing the VUS rate is a major goal in clinical genetics.
29:25And the value of these large scale projects, like a 100 kill EGP. Immense for discovery. They allow systematic searches and provide the statistical power to find these enrichments. They also fuel tool development.
29:36But as the UK biobank comparison showed, the richness and accuracy of the clinical data linked to the genomes is paramount, garbage in, garbage out still applies, even with big data. Let's circle back to that RNA validation bottleneck.
29:50It really seems key. It is. The low FBN one expression in blood is a persistent challenge. RTPCR work here with effort, but sometimes other approaches like using skin fiber blasts or perhaps even novel methods might be needed.
30:04The difficulty in getting robust validation definitely contributes to diagnostic delays and uncertainty for some families. And getting that validation is especially important because FBN one variants can be found incidentally, right?
30:16It's on the secondary findings list. That's a very important point. If you find one of these cryptic variants by chance when sequencing for something else, you need strong proof it's actually damaging before reporting it back and potentially causing unnecessary worry or interventions.
30:29RNA evidence provides that crucial functional support. Do you think the production tools will get so good eventually that we won't need as much lab work? They're constantly improving. Splice AI is a huge leap over older tools.
30:43Future AI might become even more accurate, maybe incorporating gene specific knowledge. Perhaps computational evidence could contribute more weight in classification, reducing the need for wet lab work for some variant types, but completely replacing experimental validation, especially for complex or novel splicing patterns.
31:01I think that's unlikely anytime soon, given the biological complexity. Any other interesting tidbits? They mentioned that for some genes, these deep intronic variants that cause pseudoexons seem to cluster in certain areas, perhaps because the underlying sequence is already somewhat primed to be recognized as an exxon if a mutation occurs nearby.
31:20Understanding that could refine searches too. So overall, this paper is a fantastic case study in pushing the boundaries of genetic diagnosis. It really is. It shows that for a condition we thought we understood reasonably well, like Marfan, looking deeper into the genome with the right combination of sequencing, computational analysis and lab work, can still yield crucial answers for families waiting years for a diagnosis.
31:44With complex demanding work, but the payoff and ending that diagnostic odyssey and enabling better care is just enormous. Absolutely. Providing certainty, allowing proactive family screening it can literally save lives by managing the aortic risk effectively.
32:00Which leaves us with a really provocative thought. If these cryptic variants are causing 3% of unexplained Marfentade, how many other genetic conditions currently listed as unsolved, have their answers lurking in the introns or other non-coding regions, just waiting for us to apply these kinds of comprehensive approaches?
32:17That's the big question driving a lot of genomics research right now. This study provides a compelling glimpse of what might be hiding in plain sight, if only we look in the right way. A fascinating area.
32:28This exploration was based on the open access article titled Utility of Genome sequencing and group enrichment to support splice variant interpretation in MarFan Syndrome by Susan Walker and Poleagues, published in genetics and medicine in 2025.
32:42It's available under a CCBY 4.0 license. And you can find the DOI and a link to that license right in the description for this session, so you can check out the original paper for yourselves. Thanks for joining us for Base by Base.
32:54We'll be back soon to explore another piece of complex science.