This episode examines a minigene-based functional study of 52 CHEK2 splice-site variants from the BRIDGES project, reporting widespread splice disruption, characterization of 89 transcripts, and an ACMG/AMP-informed tentative clinical classification.
0:19Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Thanks for having me back. Yeah, of course.
0:29All right, so let's just jump right in. Today's deep dive into the source material is really on a mission to solve, I think, one of the most frustrating problems in modern medicine. Oh, absolutely. It's a huge issue Right.
0:44Those genetic test results that just they leave you with more questions than answers. Yeah, which is the last thing you want from a test. Exactly. So we have this really fascinating study today that maps out exactly what happens when our cellular machinery misreads our DNA.
0:59And, you know, how researchers are building these isolated test environments to finally give patients clear answers about their breast cancer risk. It's pretty groundbreaking work, honestly. So imagine for a 2nd that you're the one taking that genetic test.
1:13You're trying to understand your personal breast cancer risk, right? You get the results back, and you're bracing yourself. You're ready for either a high risk or hopefully a low risk verdict. Right, exactly.
1:23But instead the report essentially says, um, we don't know. Yeah. It's a huge, incredibly frustrating question mark for anyone in that position. For sure. In the medical world. This is called a variant of uncertain significance or a V US for sure.
1:40Of the US, right? Yeah. So the lab found a typo in your DNA, but they simply don't have enough data to tell you if that typo is harmless or if it's, you know, going to cause cancer. And to really understand why this happens.
1:53Well, we have to look at how ourselves actually read our genetic instruction manual in the 1st place. Right, the basic biology of it. Yeah. Think of your genes like a giant recipe book. The ingredients, the actual code for making the proteins, they're all correct.
2:07Like, the words are spelled right. Okay. The paragraph breaks and the spacing between those words are shifted, just a tiny bit. Suddenly, the instructions become this completely chaotic mess. I love that analogy because, you know, you might have the right ingredients, but you're basically putting the cake in the oven before you've even mixed the batter.
2:23Exactly. It's a disaster. And the mechanism there is just completely fascinating. It revolves around a process called pre-MRNA splicing. Yeah, so before a cell can actually use the genetic code to build a protein.
2:38It has to edit that code. It acts kind of like a molecular film editor. Okay, so cutting and pasting. Exactly. The raw genetic footage has a lot of junk sequences, which we call introns. The cell has to cut those introns out and paste together the crucial instructions, the exxons to make a coherent movie.
2:55Right, so the Exxons are the actual recipe steps you need. You got it. And this cutting and pasting relies on very specific genetic signposts called splice sites. If you have a tiny typo at one of these spice sites, I mean, literally, just one single letter out of place, the cell's editing machinery gets horribly derailed.
3:14Wow, just one letter. And when that derailed process happens in a gene that's designed to protect you from tumors, the consequences are obviously severe. Their life altering, yeah. Right. So to solve this specific, we don't know problem.
3:27A massive collaborative effort was required. Today we celebrate the work of the International Research team from the Spanish National Research Council, Hospital of Clinico San Carlos, the University of Cambridge, Leiden University Medical Center, and the Bridges Project, who have advanced our understanding of genetic variant classification in breast cancer.
3:46And to really appreciate the scale of what this team accomplished, we need to focus on the specific gene they targeted, which is chi 2. Right, check 2. I've seen it described in the sources as a genome guardian.
3:59But what does it actually do on a molecular level? Like, how does it guard? So cheat to encodes a kinase protein. Think of a tiny as like a molecular switch. It's an enzyme that attaches tiny chemical tags, these phosphate groups to other proteins, which basically flips them into their active state.
4:16So GK2 essentially patrols the nucleus of yourself. Just looking for trouble. Pretty much. When your DNA suffers a double strand break, which is a catastrophic type of cellular damage where the DNA helix just snaps completely in half.
4:30That sounds bad. It's very bad. When that happens, another protein called ATM spots the damage and it flips the switch on cheek two. Okay, so Che 2 gets activated. Right. T2 then signals a whole cascade of downstream proteins to just rush in and repair the DNA.
4:44If your sheet 2 gene is broken, that repair crew never gets the call. Oh I see. And then mutations just start to accumulate wildly. Actually, the brakes are off. The skakes there are incredibly high for the patient then.
4:55Looking at large scale studies, like the Bridges project, protein truncating variants, which are basically severe mutations that cut the protein short, right? Yeah, they ruin the protein completely. Right.
5:07So those variants in Cheat 2 account for nearly 25% of all such variants across 8 major breast cancer genes. It's a huge proportion. Yeah. It is second only to the infamous BRCA2 gene in terms of prevalence.
5:20It really is a heavy hitter. Carrying a pathogenic cheat 2 variant gives a person about a 25% absolute risk of developing breast cancer by the age of 80. Wow. One in four. Yeah. And it also increases the risk for other cancers too, like prostate and colorectal cancer.
5:37The risk is substantial enough that medical guidelines actually recommend carriers begin getting annual breast MRIs starting at age 30 to 35. Which, I guess, brings us to the central roadblock of this whole deep dive.
5:49If we know cheek 2 is so critical for preventing cancer, and the screening guidelines are very clear for high risk patients. Why are over half the test results a total mystery? The big question, right?
6:01Yeah. The paper notes a staggering 54% of all cheat 2 variants, reported in the clinical database Klenvar, are classified as variants of uncertain significance. Which is just an unacceptable number. It really leaves patients and doctors completely in the dark.
6:17Why is it so hard to just sequence the DNA and, you know, know what the splicing mechanism is broken? Well, it comes down to the mechanics of how we typically test for these splicing errors in a living person.
6:28Usually researchers would take an RNA sample from the patient's blood to see if the splice transcripts are coming out normal. Okay, that makes sense. Just look at the final product. Right, but there's a catch.
6:38Humans inherit 2 copies of most genes, one from each parent. So the patient usually has one mutated copy of check 2, but also one in a perfectly healthy, wild type copy. Oh, I see. So the healthy copy is still doing its job.
6:52Exactly. And that healthy copy acts as a massive confounding factor in the lab. It's just pumping out normal RNA transcripts, which completely drowns out, or like masks the exact errors being caused by the mutated copy.
7:04Oh, wow. So you can't even see the mistakes. Yeah, it's like trying to listen to a faint static filled radio station while a perfectly clear station is broadcasting on the exact same frequency. You just can't isolate the bad signal to study it.
7:16That sounds incredibly frustrating. So if the patient's blood is too noisy. How did they isolate that bad signal? Did they, um, use CRISPR to knock out the healthy gene in a cell culture or something? I mean, that would be one way, but it's incredibly difficult to do at the scale required for this many variants.
7:34Instead, they used a really brilliant approach called a single allele strategy. Single allel. Yeah, leveraging something called minagene technology. But, uh, 1st they had to figure out which variants to even test.
7:46You can't test everything. Right. 54% of the database is 1000s of variants. Exactly. So they started with 128 unique chick T2 variants identified at the borders of introns and Exxons. They gathered these from the Bridges project.
8:00And just for contest, that's a massive database, looking at over 110,000 breast cancer cases and controls. Yeah, it's a huge data set. A data set of 110,000 people really gives you an incredible cross-section of humanity's genetic typos.
8:15It really does. So they ran those 128 variants through some advanced bioinformatic algorithms. They use a tool called max and scan and a deep learning neural network called Splice AI. Oh, Splice AI. So it's an AI that predicts the splicing outcomes.
8:29Yeah, exactly. It predicts it directly from the raw sequences. This bioinformatic filtering allowed them to basically whittle the list down to 52 highly suspicious variants. Okay, so the prime suspects.
8:40Right. The ones strongly predicted to wreck the splicing process. Okay, so they have their 52 prime suspects. Now comes the minig part. If they aren't editing a live cells native genome. What exactly is a minogene?
8:52It is essentially a synthetic stripped down version of the gene. They engineered three custom minogenes that together span all 15 exxons of the CHK2 genes. Okay. They place these genetic sequences into a specialized delivery vehicle called a PSAD vector, which stands for splicing and disease, by the way.
9:13Oh that's convenient. Yeah. And then they introduce these vectors into MCF 7 cells. Wait, let me pause you there. YMCF 7 cells. Why not just put this synthetic gene into yeast or, you know, some basic easy-to-grow lab, so?
9:27Because the cellular environment matters immensely. You see, MCF 7 cells are a line of human breast adenocarcinoma cells. Oh, so they're breast cancer cell. Exactly. Splicing isn't just a passive process.
9:38It relies on a whole cocktail of specific proteins and factors floating around in the nucleus. Breast tissue has a distinct profile of these splicing proteins compared to, say, liver tissue or yeast. Oh, that makes perfect sense.
9:49By using breast cancer cells, they're studying the variant in the exact cellular environment where the disease actually originates. You got it. It's the most accurate context you can get outside of the human body.
10:02So it's like instead of trying to read the blurry original recipe in the patient's busy kitchen, they built this pristine, isolated test kitchen. About the test kitchen analogy, yeah. They put the mini gene into the breast cancer cell and they introduce just one genetic typo at a time.
10:18It's a test kitchen with a massive catch, though. Because there is no healthy gene copy in this synthetic setup, whatever RNA transcripts are produced comes solely from the minogene. Right, no background noise.
10:29Exactly. This allows the researchers to detect even the tiniest amounts of mutant transcripts with absolute clarity. They track every single time the cell drops a bowl or, you know, reads the wrong line of the recipe.
10:41Okay, so they ran these 52 prime suspect variants through the midigene test kitchen. What did the cells actually bake? Well, the failure rate was staggering, honestly. An overwhelming 46 out of the 52 variants that's 88.5% actively impaired the splicing process.
10:58wow Yeah. And of those, 34 variants completely destroyed the process entirely. Like they left 0 trace of the normal full length transcript. That's wild. And the paper notes that the sheer variety of the mistakes the cell made was mind boggling.
11:14I really need to unpack this. They tested 52 tichos, but they annotated 89 completely different flawed transcripts resulting from them. Yeah, it's a lot of errors. How does one single typo cause the cell to create up to, like, 11 different wrong recipes for a single variant.
11:31It feels like the cell is actively trying to invent new ways to fail. Well, the cell is panicking, basically. The mechanism here is what we call alternative splice and chaos. Spice and chaos. sounds accurate.
11:41Yeah, so remember the spice of some, the molecular machine responsible for editing the genetic film. When a normal splice site is broken by a mutation, the spicy sum doesn't just stop and give up, it frantically grabs at other nearby sequences that look even vaguely like a splice site.
11:55Oh, wow, just desperate. Very desperate. We called these cryptic splice sites because the primary binding site is ruined. The splay system attaches to these weaker, secondary sites in a totally unpredictable manner, just churning out a multitude of bizarre combinations.
12:11So it was just guessing. Pretty much. Sometimes it skips a single Exxon, sometimes it skips multiple Exxons at once. Sometimes it leaves chunks of junk intronic DNA right in the middle of the instructions.
12:23Imagine you're the patient holding that test result. Knowing your cells might be making 89 different corrupted versions of a protein that's meant to protect you. It's a sobering thought. And the end result of these broken recipes is usually catastrophic for the protein itself, right?
12:39Out of those 89 flawed transcripts, I saw that 59 of them were predicted to introduce a premature termination code on. Yeah, which is exactly what it sounds like. It's a genetic command that basically screams, stop cooking immediately, right in the middle of the sequence.
12:53Wow. So the protein is cut short, completely nonfunctional, and the guardian of the genome is totally disarmed. Now, that frantic grasping by the spicy sum, leads us to, I think, one of the most surprising discoveries of the entire study.
13:06They looked really closely at one specific variant, designated C.6842AG. Right. fascinating. Yeah. When they ran this variant through the Minogene. It activated a bizarre, non-canonical TG acceptor site.
13:21Now, TG Acceptor sites are incredibly rare. The paper noted they make up only about 0.02% of all humans play sites. Yeah, they are practically non-existent. Right. How is the cell pulling a nearly non-existent punctuation mark out of thin air?
13:36So the standard sequence for an acceptor site almost always ends in the letters AG. This specific typo forced the splice system to recognize a TG sequence instead. And to figure out how the cell was pulling this off, the researchers played this meticulous game of molecular elimination.
13:51Okay, what did they do? They started engineering microdilutions into Exxon 6 of the Minogene. Wait, so they were deleting tiny chunks of the code around the typo to see when the rare TG site would stop working?
14:02Exactly. They realized this rare splice site wasn't acting alone. It was being propped up by architectural helper proteins, essentially. molecular scaffolding sitting further down the DNA line. Oh that's so clever.
14:15They removed a 37 nucleotide chunk here, a 14 nucleotide chunk there, and they found hidden sequences called exonic splicing enhancers. When they deleted the regions containing the scaffolding, the whole rare TG splicing event just collapsed.
14:31The cells stopped using it entirely. Wow. So the surrounding structural context of the gene was secretly dictating how the error manifested. That really highlights just how exquisitely complex pre-MRNA splicing really is.
14:45It's not just a linear code, it's a whole 3D structural environment. It demonstrates that you can't just look at the typo in isolation. You have to look at the entire neighborhood around the typo to understand how the cell is going to react to it.
14:57This biology is incredible. It really is, but I want to connect this back to the patient sitting in the doctor's office. Right, the clinical application. Yeah. How does mapping molecular chaos in a minigene actually help the person holding that we don't know test result?
15:10Well, this is where the bench science translates directly to clinical care. The team didn't just stop at describing the 89 weird transcripts. They took all of this pristine minagene red out data and fed it into a highly rigorous clinical framework provided by the ACMG and AMP.
15:27Right. those are the guidelines in the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. They use a complex point-based system, right? They do. It's very systematic Positive points for pathogenic evidence, negative points for benign evidence.
15:41A score of +10 or more means a variant is definitely pathogenic. +6 to 9 is likely pathogenic. Okay. So the researchers integrated their minogene results into this framework, assigning specific strength codes based on the severity of the splicing errors they observed in the lab.
15:57And by doing this, they successfully provided tentative clinical classifications for 32 of the 52 variants. Wow, 32 of them. Yeah, they classified 27 as pathogenic or likely pathogenic, and 5 as likely benign.
16:10That is a direct, actionable answer. And crucially, they went into the Clinvar database and looked at variants that previously had conflicting data from multiple submitters. Using this new evidence, they were able to reclassify 8 of those conflicting variants.
16:26That's huge deal. Yeah, for example, they moved 3 variants of uncertain significance into the pathogenic or likely pathogenic category. And we really cannot overstate how life altering that changes. It means the families carrying those three specific variants can now justify and receive early MRI screening and preventative care.
16:46Which is amazing. Catching breast cancer early saves lives. And conversely, moving 2 variants from uncertain to likely benign saves those families from years of unnecessary anxiety and invasive unneeded screening.
16:57I hear that. And it's incredible that they gave definitive answers to those families. But, you know, I have to push back a little on the math here. Out of the 52 variants, they tested, 20 variants that's 38% are still stuck as BUSs.
17:10They remained in US limbo, even after all this exhaustive miniagene testing. Doesn't that mean this minaging tool has a pretty hard ceiling? Like it didn't solve everything? It definitely has a ceiling, yeah.
17:21And it's a critical limitation we absolutely have to acknowledge. The minigene essay is incredibly powerful, but it has boundaries. Those 20 variants stayed in the uncertain category for a couple of reasons.
17:32First, as we discussed, splicing is incredibly nuanced. Several of these variants exhibited what we call a leaky splicing. Leaky splazing. That's where the splice site is damaged, but not utterly destroyed, right?
17:44So the cell produces some normal full-length transcripts alongside all the broken ones. Exactly. Because chemistry inside a cell is probabilistic. It's not absolute. Sometimes the splices some still manages to bind to the damaged site properly.
17:57Oh I see. If a variant is leaky and it's still producing, say, 20 or 30% of the normal protein. The question becomes, is that enough to prevent cancer? Or is the drop in protein levels sufficient to increase risk?
18:10The minigene alone can't answer that clinical threshold question? You still need large scale family data to know if that specific protein drop leads to tumors in real life? That makes total sense. And there's also the issue of genomic context, right?
18:24You mentioned earlier that the menogene is a synthetic stripped down version of the gene. Right. A minogene lacks the total massive genomic context of the native DNA folded up inside a human chromosome.
18:37While the researchers went to great lengths to include the natural flanking sequences for each target Exxon, sometimes regulatory elements sitting incredibly far away from the gene can actually influence splicing.
18:49Right, like those enhancer scaffolding sequences they found in Exon 6, but potentially on a much grander scale across the whole chromosome. Exactly. So the lack of full genomic context is an inherent limitation of the tool.
19:00However, even with that ceiling, the value of what was achieved here is immense. Oh, absolutely. Yes, 20 variants remain uncertain, but 32 are no longer a mystery. Moving even 8 variants at a Clinvar conflict status is a massive clinical win that directly alters patient management.
19:17It really shifts the entire paradigm from passively waiting for statistical family data to actively interrogating the molecular mechanism of the disease. Yeah, that's a really powerful way to look at it.
19:27So, to bring this all together, by building isolated genetic test kitchens called minogenes, researchers systematically map the chaotic splicing effects of 52 GK2 variants, translating genetic typos into actionable breast cancer risk classifications.
19:42beautifully summarized. Thanks. This approach cuts straight through the ambiguity of uncertain test results, directly improving patient screening, and ultimately saving lives. It is a profound demonstration of how fundamental molecular biology, literally understanding exactly how a splices some panics and makes a mistake can reach straight out of the lab and protect a patient's health.
20:02It certainly makes you view your own genetic code with a lot more respect for the delicate, high stakes editing process happening inside every single one of your souls right now. And I think it leaves us with a really powerful question.
20:14What does this mean for the 1000s of other mysterious genetic variants sitting in databases today waiting to be understood? A lot to think about. It really is. This episode was based on an open access article under the CCBY4.0 license.
20:29You can find a direct link to the paper and the license in our episode description. If you enjoy this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
20:41Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.
21:12Be on home on the bench, clean hands. Steady gaze, one long thread of lettuce. Ooh, the microscope pays from five prime to three, we don't cut, we don't splice. We ride the whole waveform, pay the full price, long reef, so burns, lying up in time, hit rooms in the entrance, footprints in the rhyme.
21:37The pattern breaks the silence. don't look away. We phase it to the truth, till the shadows obey, full, full, light. Let that apple type speak 26 kilo bases running strong, honestly. When parents try to vanish, we can still see.
21:56Faced in the dark now is clear to me. Phased in the dark, now is click... I repeat in the first, a Pintron. Counting like a drum, ETR, modies folding where the signals come from a missing stretch aside, a switching enhancer, torn recombinant, lines crossing, chime reborn.
22:31So if the blood types burn, the match feels like the unseen structure in the every skin, not a guess, not a maybe, just a better way to read what was written, got what we said. All in full light. Let the hypotypes be 26 kilo bases, running strong, running sleep.
22:54Big deletion, stranger, combined, set them free. Phased in the dark now is clear to me.