A concise walkthrough of data-driven heuristics and a splicing checklist derived from large-scale exon, branchpoint, and experimentally validated variant analyses to improve interpretation of splice-altering variants.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and read us in your podcast app. Yeah, thanks for joining us for another deep dive.
0:10So, imagine for a 2nd that you are a clinical geneticist, right? Okay, sitting the scene. Yeah, so a family is sitting in your office and they are just desperately seeking an answer for their child's undiagnosed condition.
0:21You decide to sequence the patient's genome. Which is standard practice now, but it's a massive undertaking. Exactly. You are scanning through like 3000000000 letters of DNA. And buried within that massive code, you find a single, tiny typo.
0:35Just one altered DNA letter out of billions. Right. And now you are faced with this massive high stakes question. Does that one microscopic typo cause a devastating disease or is it, you know, just harmless?
0:50Is it just completely normal human variation? Yeah, it really is the ultimate needle in a haystack problem. It has to be incredibly stressful. Oh, absolutely. And what is so profoundly critical to understand here is that when we do find these disease causing genetic variants.
1:05Um, between 10 and 30% of them don't actually break the core blueprint of the gene itself. Wait, really? So what do they break? Well, they break the cellular editing process that translates that blueprint.
1:17It's a process known as MRNA splicing. Okay, let's unpack that editing process for a moment. Think of your DNA like a massive unedited manuscript for a highly complex technical manual. like that analogy.
1:32Yes, you've got the crucial chapters that actually tell the cell how to function, right? Those are called exxons. But sitting right between those essential chapters, you have 100s of pages of behind the scenes formatting code.
1:42All the directors commentary and the fluff. Exactly. Biologists call these intervening non-coding regions introns. So MRNA splicing is like a master editor coming in, carefully cutting out all that unnecessary formatting code.
1:55The introns. Right. cutting out the introns and stitching the remaining chapters together so the final manual flows perfectly. But if a genetic typo removes a crucial punctuation mark, while that cellular editor might get terribly confused.
2:07Yeah, it might accidentally delete an entire vital chapter or, you know, leave pages of raw formatting code right in the middle of the book. Precisely. And the physical machinery performing this editorial work is called the Splicism.
2:20But predicting exactly which single letter typo will derail the splice of sum is incredibly difficult. So looking at a sequencing report and trying to predict splicing outcomes has basically felt a bit like a guessing game.
2:33It has for a long time. But how could changing the way we interpret these minute genetic errors transform clinical diagnoses from that guessing game into an exact predictive science? Well, today we celebrate the work of Patricia Sullivan, Mark Calley, and the research team at the Children's Cancer Institute, and UNSW Sydney, who have advanced our understanding of splice altering variants.
2:56And their work was published in the American Journal of Human Genetics in 2025. It's a fascinating deep dive. It really is. So to appreciate the magnitude of what this team has accomplished, we have to look closely at the clinical problem they're trying to solve.
3:10Right. Splicing, as you mentioned, is orchestrated by the Splice SM. But the spice system isn't just some abstract concept. It is a massive physical complex made of RNA and proteins. It's basically a bulky molecular machine.
3:24Exactly. And for this machine to know exactly where to land, where to cut, and where to paste. It looks for highly specific sequence motifs scattered throughout the RNA string. Right. There is this whole specialized ecosystem of landing pads essentially.
3:39Yes very specific ones. Like at the beginning of the non-coding region, you have the donor splice site. Also called the 5 prime splice site. Right. And then at the end of that region, you have the acceptor splice site.
3:49And it gets even more complex because hidden deeper inside the non-coding sequence. You have a branch point sequence. And right next to that, something called a polypyrimidine tract. Okay, so if you get a genetic mutation in any of these critical zones, you alter the physical landing pads.
4:06Exactly. You create what is called a splice altering variant or an SAV. And this brings us directly to the core tension in clinical genetics today, doesn't it? It does. When geneticists are trying to interpret these variants for a patient relying on standard evaluation frameworks like the ACMG and AMP guidelines, they lean heavily on artificial intelligence.
4:26Like in silico computational prediction tools. You hear names like splice AI or pangolin. Right. And to be fair, those tools are incredibly powerful at what they do. But they're essentially black box models.
4:37Yes, that is the big issue. They are trained by analyzing 1000000s of naturally occurring healthy RNA transcript sequences. So they aren't explicitly trained on experimentally validated real-world mutations.
4:51No, they aren't. So they suffer from this massive structural blind spot. They can look at a patient's DNA and predicts that a splay site is disrupted by a mutation, but... But they completely fail to predict the actual outcome on the transcript.
5:03Yeah, exactly. Yeah. Which leaves the clinician in an incredibly frustrating position. The AI raises a red flag and says, warning, this site is broken. But it cannot tell the doctor what the cell actually does in response.
5:16Does it skip the whole Exxon? Does it leave jump DNA in? The AI simply doesn't know. So the clinician doesn't have the definitive evidence required to confidently diagnose the patient. Okay, I have to push back on this a bit, though.
5:29If you're listening right now and thinking about the state of AI today, you might be wondering, if AI can drive a car through a busy city, why does it struggle to just read a linear string of DNA? It's very fair question.
5:42I mean, let's look at the specific PKD one gene example from the source material. The AI correctly flagged that a mutation broke a specific splice site, right? It did. And it even identified that there were 3 alternative backup sites nearby.
5:57Right, but then the AI completely failed to predict which of those 3 backup sites the splice of stone would actually use to try and fix the error. Why is a multimillion dollar computational model so bad at figuring that out?
6:09Because it is a fundamental limitation of how the AI views the world. These black box models look at DNA as a flat linear string of text. Oh, but it's not flat. Not at all. The spicy sum is a physical three-dimensional machine.
6:24It has strict spatial and mechanical constraints. So in that PKD one example, the physical folding of the RNA dictates the outcome. Exactly. The spatial rules of assembly mean that only one of those 3 alternative backup sites could possibly be used.
6:39The other 2 were physically inaccessible. Wow, so the AI just didn't know the physical rules of assembly. Right. It just looks at text patterns. Because black box AI lacks this transparent biological reasoning.
6:51The researchers realized they couldn't just build another algorithm. They needed to build a comprehensive evidence-based framework from the ground up. Yes, based on physical biological rules. So instead of training a neural network, they decided to build a transparent biological checklist.
7:04But to do that, they 1st had to define the physical boundaries of what is normal. Right. Right. Step one. They analyzed over 202,000 canonical protein coding exxons. Which is a massive data set, and alongside that, 19,000 experimentally validated branch points.
7:22Specifically focusing on the major U2 splice system, right? The primary editing machine in ourselves. Exactly. By analyzing this data, they basically reverse engineered the absolute minimal sequence and spacing requirements for splicing to happen in a healthy cell.
7:37So that established the baseline. But knowing the rules for healthy genes is only half the battle. Right. Right. They had to see what physically happens when those rules get broken. So for step two, they move to empirical variant analysis.
7:48They pulled over 11,800 experimentally validated single nucleotide variants from Splice VardDB. And we really should clarify what experimentally validated means here because this isn't computer guesswork.
8:01Oh, right. This is data where a scientist actually grew cells in a lab, introduced the exact mutation, and physically measured the resulting RNA to prove what happened. Exactly. It is ground true biological reality.
8:14And this database included nearly 9000 variants proven to actively alter splicing, and nearly 3000 that were proven to be perfectly normal. Okay, so step three. What do you do with all this ground truth data?
8:27You start creating heuristics. Biological rules of thumb. So they grouped all these variants based on their exact physical location and created decision trees. But they instituted this brilliant, strict quality control rule.
8:40They required a minimum of 10 functionally validated variants per heuristic. Oh, to avoid overfitting the data. Precisely. If they didn't have at least 10 real world lab proven examples of mutation at a specific position.
8:53They just refuse to make a rule for it. That's smart. And that allowed them to establish a powerful new metric called Splicogenicity. Yes, Swissogenicity is the exact proportion of variants at a specific location that actually end up affecting splicing.
9:06And that completely shifts the paradigm. It moves us away from those opaque binary yes or no predictions. It provides a nuanced, context driven probability matrix. Right, instead of a computer just flashing a warning light, this framework tells a clinician, hey, a mutation at this exact physical position has a 95% probability of destroying the splice site.
9:27It forces biological reality back into the diagnostic process. Okay, so what are the baseline rules they actually found? Well, they found that 95.9% of normal major splice of some introns pass their newly defined checklist.
9:41And one of the most fascinating constraints on that checklist is the length. The minimum intron link is 80 nucleotides. Yeah, the non-coding region has to be at least 80 letters long. Because the splice system is a bulky bulldozer.
9:54If the track is shorter than 80 letters, the machinery literally doesn't have enough physical room to assemble. Exactly. It would just crash into itself. You also need a minimum of 9 pyramidines, which are the letters C and T in that polypyramidine tract.
10:07Right, as a chemical landing pad. But perhaps most critically, they identified the AG exclusion zone. Yes, this is a very specific region between positions -13 and -6 right before the end of the intron.
10:20And in healthy genes, the letter combination AG is almost entirely depleted in this zone. Because the splice of some uses a scanning mechanism to find the end of the intron, looking for an AG to make its final cut.
10:31Oh, so if there are stray AG sequences sitting in that exclusion zone, the spice of some might grab the wrong one and misfire. Exactly. Evolution essentially swept that zone clean. Wow. Okay, so we have the rules of the healthy road, but when you look at the mutations that break these rules, The discoveries are wild.
10:50Oh, absolutely. Let's look at the donor site at the very start of the non-coding region. The numbers here are crazy. 93.8% of validated variants at the donor site ultra-splacing. It is incredibly fragile.
11:02And the specific details matter, like the last base of the actual Exxon, the E -one position. The splice also has a massive rigid preference for which letter it wants there. G is highly preferred. A is acceptable.
11:15C is problematic, and T is the absolute worst. So if a mutation introduces a T instead of healthy G, it just fundamentally breaks the physical hydrogen bonds. Yes. And at the 5th intronic base, the +5 position.
11:27The outcome heavily depends on whether the original reference nucleotide was a G. If it wasn't a G to begin with, the spliciosum doesn't really care about mutations there. Right. But if there was a G and it changes, it almost always disrupts the process.
11:42But then you look at the acceptor site at the other end, and the dynamics are completely different. Only 61.one% of variants here are splice altering. Yeah, it's much more resilient to single letter changes.
11:53Why is that? Because the acceptor site requires a much broader sequence context. It relies on that scanning mechanism we talked about. Right. Right. Sliding along the RNA string, looking for the 1st AG.
12:05Exactly. Because of this, it is highly sensitive to new accidental AG letters popping up upstream in that exclusion zone. So if a mutation introduces a false stop sign in the cleared out zone, the scanning splice system grabs it too early and derails.
12:21Precisely. Which brings us to the ultimate question. What does this derailment actually look like for the RNA? We know the AI couldn't quantify the chaos, but this team did. They did. And the magnitude of the disruption is staggering.
12:34In 51.3% of the cases, the result is Xon skipping. Over half the time. So the splice zone gets so confused. It just abandons the entire chapter. The whole Exon is deleted, usually resulting in a non-functional protein.
12:47And then in 36.3% of cases, you get X on truncation. Right. The master editor misinterprets the signals and physically cuts the chapter right in half. Leaving out massive chunks of essential information.
13:00And then you have intron retention, occurring 17.4% of the time. This is where it fails to recognize the non-coding region at all, and just leaves pages of raw formatting code in the final transcript. And the paper noted that donor site damages the primary culprit for full intron retention, right?
13:16It is, yeah. And finally, 16.2% result in Exxon extension, where it cuts too late and attaches junk code to the end of the chapter. It is a catalog of genetic disasters. It really is. But hold on, I have to challenge the rigidness of these rules for a second.
13:30Go for it. You and I are sitting here discussing exact percentages and strict nucleotide preference orders. But biology is notoriously messy. Oh, it's incredibly messy. Right. So what happens when a patient has a variant that is buried deep, deep in the junk DNA, far away from these neat splice sites, does this checklist just fall apart?
13:50That is an excellent point. And you were talking about pseudo-exons. Pseudo-exxon. Yeah, sometimes a mutation buried deep inside the vast non-coding region will change just one single letter. And that accidentally creates a sequence that mimics a canonical spice site.
14:05So it essentially creates a mirage, a false start signal in the middle of nowhere. Exactly. It tricks the spicy sum into including a random piece of junk DNA into the final transcript, destroying the resulting protein.
14:16So did the researcher's newly built checklist actually flag these deep space anomalies? We did. The team's checklist successfully flagged the cryptic sites for 82% of confirmed pseudoexons. Wait, really?
14:2982%? Yes. It proved that the minimal physical spacing requirements are actively used by the splicosum, even when it's being tricked in the middle of the junk DNA. That is amazing. So if we take a step back from the molecular machinery, what does this all mean for the clinician and the patient?
14:47Well, this framework bridges the gap between raw computational prediction and actual clinical variant evaluation. It gives geneticists a transparent, reproducible checklist. Yes. When evaluating a variant using the ACMG and AMP guidelines, they no longer have to blindly trust a black box.
15:04They can look at the checklist and interpret exactly why an anomaly is pathogenic. They can actually document the physical mechanism. This mutation introduces an AG into the exclusion zone, causing Exxon skipping.
15:17Exactly. It provides the concrete evidence required for a diagnosis. But as always, we need to note the study's boundaries. Where does this framework fall short? It's crucial to acknowledge the limitations?
15:28This framework doesn't account for splicing regulatory elements are SREs. Right. Enhancers and silencers that depend on fluctuating protein abundance. Yes. They give the FGB gene as an example. A variant there, simultaneously damage to silencer, and strengthened an enhancer, causing a pseudoaxon.
15:47And the core mechanical checklist struggled to predict that because SREs introduce such a complex regulatory layer. Precisely. And because of that, the model is effectively tissue independent. Meaning it misses tissue specific edge cases, since different tissues have different regulatory proteins active.
16:04That's right. But even with those limitations, the future horizons are vast. The ultimate goal isn't just a manual checklist. Right. Integrating these biological heuristics into high throughput bioinformatic tools.
16:17Making future AI predictions, gyologically sound and totally transparent. Because knowledge is only as good as our ability to apply it. So let's summarize exactly what we've unpacked today. good. By extracting hard biological rules from 1000s of experimentally validated genetic variants.
16:34Researchers have created a transparent data-driven framework for predicting splice altering variants. This moves the field away from opaque AI predictions, and toward a comprehensible checklist that accurately anticipates how a DNA mutation will scramble an RNA transcript.
16:52It is a profound shift in how we approach diagnosis. What does this mean for the future of personalized medicine when we can finally explain the why behind every genetic title? That is something we will all be watching closely.
17:04This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
17:17If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
17:26Thanks for listening and join us next time as we explore more science based by base.