Pepe et al. show that annotating cancer mutations to the transcripts actually expressed in tumors uncovers previously overlooked non-coding promoter mutations in melanoma. Using TCGA mutation calls, RNA-seq, and an automated Salmon+VEP pipeline, they reclassify multiple hotspots and validate functional effects for IRF3/BCL2L12 and KNSTRN promoter mutations with CRISPR-Cas9 and reporter assays.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So what if I told you that the massive databases oncologists use to treat cancer every single day might be reading, you know, the completely wrong instruction manual.
0:19Right. It's a pretty terrifying thought, honestly. Yeah, totally. Because when we look at our own DNA, we generally think of it as this ultimate flawless master document for our bodies. But like any massive text, it accumulates typos over time.
0:35Exactly. Yeah. And most of those genetic typos are completely harmless. In biology, we call them well, silent or synonymous mutations. Right. So imagine you're reading a recipe book, right? And the recipe for a chocolate cake has a typo that spells stir with 2 R's.
0:50Oh, I like that. Stir the batter. Yeah, exactly. You as the chef, you realize what it means. You just completely ignore the typo. You still stir the batter, the cake comes out exactly the same, and literally nobody cares.
1:00Nobody cares at all. What's fascinating here is how deeply ingrained this assumption is in the medical and scientific community. We have historically brushed right past these synonymous mutations. Because they're just, like, background noise.
1:14Yeah, exactly. Computational biologists have primarily used them just to calculate that background noise in a genome. Wow, really? Yeah. But what if we've been looking at the wrong edition of the recipe book this entire time?
1:28Wait, meaning what? Meaning, what if in the specific edition of the book that the cancer cell is actually using that typo doesn't say stir instead of stir, what if it actually changes a crucial ingredient?
1:40Oh man, so you're putting in like salt instead of sugar? Uncovering that wrong instruction manual? I mean, that would require a massive shift in how we look at cancer genome. It absolutely does. It changes everything.
1:51Well, today we celebrate the work of Daniel Pepe, Xander Jansen's came to Kersmaker and their team, who have advanced our understanding of cancer genomics by showing us. We've been mislabeling mutations.
2:03Yes, and it's such an important paper. It really is. We are diving into their massive study titled reannotation of cancer mutations based on expressed RNA transcripts reveals functional non-coding mutations in melanoma.
2:17It's quite a mouthful, but the work is brilliant. Yeah, it was published in the American Journal of Human Genetics on June 5, 2025, and it's this huge collaborative effort from KU Luvin and a bunch of other collaborating institutions.
2:30And, you know, the weight of what this team accomplished. It really cannot be overstated. didn't just find a new mutation. Right, it's bigger than that. Way bigger. They challenge the foundational assumptions of how these massive genomic databases like see bioportal actually categorize cancer data.
2:46Okay, let's unpack this. Because if these mutations truly don't change the proteins, amino acids, like if the cake is still made of the exact same ingredients, why did anyone even suspect they might matter in cancer pathogenesis?
2:58That is a great question. And it's the core of the biological mystery here, because even though the final amino acid sequence of the protein is identical, the journey to make that protein can be heavily disrupted.
3:09Disrupted how? Like in the factory. Exactly. A so-called silent mutation can actually alter how the RNA is spliced together, or it can modify how stable the messenger RNA is. Oh, I see. So if the MRNA is unstable.
3:22The blueprint basically degrades before the cell can even finish reading it. Right. And it can also change the translation speed. Sometimes a synonymous mutation forces the cell to use a really rare transfer RNA.
3:35Like a bottleneck on an assembly line. Exactly like a bottleneck. It slows down the production, which can cause the protein to fold incorrectly as it's being made. Yeah, we've actually seen this in key cancer genes before, like TP 53.
3:47But here is the real issue. The databases annotate these mutations against a reference transcript. reference transcript, meaning like a default template. Yes, a generic default template of a gene, but different tissues, and especially tumors, they express different versions or transcripts of those genes.
4:05Oh, okay, so if you judge a mutation by the reference transcript instead of the express transcript, you misclassify it. Exactly. You're looking at the wrong map. So the team realized they needed a completely new way to map these genetic errors.
4:18Right. They couldn't just look at the DNA on its own anymore. So it's like police using a heat map to see exactly where burglars keep breaking into a house. But instead of looking at a generic neighborhood blueprint.
4:28They cross referenced it with the actual blueprint of the specific house to see why the burglars chose that spot. That is a perfect analogy. If we connect this to the bigger picture. The integration of RNA sequencing is the critical innovation here.
4:43Because DNA shows the blueprint, but RNA shows what the cell is actually building. Yes, exactly. They integrated DNA sequencing with RNA sequencing, and they pulled data from the cancer genome atlas TCGA and cosmic across 17 different tumor types.
5:0017. That's a massive amount of data to filter through. It is. So they used a unique mathematical approach. They utilized 4 different algorithms, hotspot 3, hotspot 12, the concentration method and the entropy method.
5:12Okay, and what were those algorithms actually looking for? They were looking for areas where mutations heavily clustered together because that signals strong, positive evolutionary selection by the cancer.
5:22Oh, right. Like the cancer is intentionally keeping those mutations because it helps it survive. Exactly. And crucially, they utilized automated computational tools like salmon and the ensemble variant effect predictor or VEP.
5:36To map those clusters against what was actually being expressed. Yes, to map them against the transcripts actually expressed in the tumors. And then they validated all of this in the lab using highly accurate direct CDNA, long read manopore sequencing.
5:50Long read nanopore. That sounds intense. It's amazing technology. They used it in a lab model of Mel ST cells, which are melanoma models to prove their findings were real. Okay, so they do all this mapping, they look at the active work orders.
6:03Here's where it gets really interesting. Out of all 17 tumor types. Melanoma had the highest number of significant mutation clusters. It did. And the bombshell statistic they found. Yeah, this blew my mind.
6:1522%. 11 out of 50 of these mutation clusters and melanoma were misannotated. Yep, almost a quarter. They weren't synonymous or misence coding mutations at all. They were functional non-coding promoter mutations.
6:28Exactly. Because the tumor was using a different transcript. Those mutations fell outside the coding region entirely. They were in the promoter. Basically, the volume dial for the gene. Yes. The volume dial is the perfect way to think about it.
6:42And the best example they highlighted involves the shared promoter region for 2 specific genes, BCL2L12 and IRF 3. Okay, so what happened there? Well, a mutation there was previously labeled as a harmless synonymous change in BCL 2L 12.
6:57Just a typo. Right. But it actually sits right in the shared promoter for both of those genes. So they used CRISPR Cast 9 to introduce the single nucleotide change into primary melanocyte models, those MLST cells.
7:10Just to see what one tiny typo would do. Just that one letter change. And they prove that it down regulates IRF 3. It down regulates BCL 2L12, and this is the kicker, the major tumor suppressor protein, TP 53.
7:22So these harmless biological typos are actually actively sabotaging the immune response and shutting down tumor suppressors like TP 53. Yes, exactly. They're actively sabotaging the cell. But how does changing one letter in the non-coding region do all that?
7:36Well, that mutation disrupts the binding sites for ETS family transcription factors. Wait, transcription factors, those are like the keys that turn jeans on and off, right? Exactly. So by disrupting that binding site, the mutation is essentially breaking the on switch for these genes.
7:51Wow. That is terrifyingly efficient to the cancer. And they looked at real clinical data for this too, didn't they? They did. In melanoma patients, this specific mutation is associated with a 19% higher rate of non-response or stable disease when treated with immune checkpoint therapy.
8:0719%. That's a huge clinical impact for something we used to call background noise. It really is. And it wasn't just BCL 2L 12. They briefly mentioned similar reannotations found in the KNSTRN and SLC 27 A5 genes.
8:21Right. Like the KSTRN mutation also disrupts an ETS transcription factor binding site. Exactly. So we're seeing this pattern where functional promoter mutations are hiding in plain sight. So what does this all mean?
8:33If I'm an oncologist or a researcher, pulling data from one of these databases today. Does this mean I am essentially flying blind to 22% of these specific types of mutations in melanoma? This raises an important question and the answer is, unfortunately, yes.
8:47Wow. Without this dual DNA and RNA approach, researchers might design the wrong experiments entirely, they might treat a non-coding promoter mutation as if it were a protein coding synonymous one. Which means they're trying to fix the wrong problem.
9:01Exactly. But the good news is, the researchers proved that their simple automated salmon and VEP pipeline operates with 90% accuracy. 90% accuracy in identifying the correct express transcript and reannotating the mutations.
9:17That's incredible. It is. The major implication here is that cancer genomic databases must stop relying solely on reference transcripts. They have to start integrating RNA sec data to reflect what tumors actually express.
9:29Definitely. Though we should know the study's limitations, right? Like the number of patient samples with these specific mutations was pretty small. Yeah, that's true. And that does limit the statistical power of the immunotherapy findings we just talked about.
9:41So we need bigger cohorts to be absolutely sure. Right. Also, their mathematical algorithms, like the concentration method. They're very strict. So less frequent driver mutations might still be missed by their pipeline.
9:53So there could be even more of these hiding in the data. Very likely, yes. Well, to wrap this all up by integrating DNA and RNA sequencing to observe what tumors actually express. Researchers discovered that over a 5th of supposed silent or mis sense mutations in melanoma are actually functional, non-coding promoter mutations driving the disease.
10:15This exposes a massive blind spot in current cancer databases and prove the critical need to update our genomic reference maps. What does this mean for all the other types of cancer where we might still be relying on the wrong reference manual?
10:27This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
10:42If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
10:52Thanks for listening, and join us next time as we explore more science based by base.