RNA-seq in human lung (LTRC) and blood (COPDGene) reveals COPD-associated variants colocalize with splicing QTLs, implicating FBXO38 cryptic exon NMD and BTC exon4 isoform shifts.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So why do some lifelong, heavy smokers escape with, you know, totally healthy lungs, while others develop debilitating chronic obstructive pulmonary disease or COPD?
0:19Yeah, it's such a massive question in pulmonary medicine. Right. I mean, we know genetics play a massive role, and scientists have mapped out the genetic neighborhoods linked to the disease, but knowing where a genetic variant lives doesn't tell us what it's actually doing to the body.
0:33Exactly. We have the location, but not the mechanism. Yeah. So how could mapping the hidden splicing of our RNA finally reveal the true culprits behind this disease? And what really happens when our genetic code gets mistranslated?
0:45It is really the core challenge of modern complex trade genetics right now. Moving from a statistical locust to a functional biological mechanism is just incredibly difficult. Today we celebrate the work of researchers from Brigham and Women's Hospital, UC Santa Cruz, and the NHLBI Top Med program, who have advanced our understanding of the genetic mechanisms underlying COPD.
1:08Okay, let's unpack this. COPD causes irreversible airflow obstruction, and while previous genome wide association studies, or GOAs, found 82 specific genomic regions linked to COPD risk. The causal mechanisms for most remain a total mystery.
1:26They really do. And, um, the standard approach to bridging that gap has always been expression, quantitative trait low si. Which we usually just call EQTL analysis, right? Right, EQTLs. You basically look for variants that change the overall amount of a specific transcript.
1:40The assumption being that a disease associated variant might simply, you know, turn up or turn down a critical gene. Like a volume dial. Exactly, like a volume dial. But using that GWA's data is like knowing a criminal lives in a specific zip code, but not knowing which house they are in or what they look like, right?
1:56The EKTL approach seems to have kind of hit a ceiling. Oh, absolutely. It's hit a massive ceiling. If we look at the data from the last few years, EQTLs really only explain a fraction of the GW signals for complex respiratory traits.
2:09So the question is, where is all that missing functional variation? Right. If it's not the volume, what is it? Well, the hypothesis driving this research, is the variation isn't in the overall expression level, but in the isoform structure itself.
2:22And that brings us to splicing quantitative trait loSi or skewTLs. Okay, SkewTLs. So we're talking about how the pre-Messenger RNA is actually processed and assembled. Yes. Alternate splicing dictates which Exxons are kept and which introns get cut out.
2:37So an HQTL is a genetic variant that, uh, statistically correlates with a change in that exact splicing pattern. Oh, I see. So the gene might be transcribed at the exact same overall volume, but the actual sequence, the actual recipe of the resulting mature MRNA is fundamentally different.
2:53Precisely. You're changing the recipe, not the volume. That makes a lot of intuitive sense. But, I mean, applying this to a complex disease like COPD requires massive scale, doesn't it? Looking at RNA from living human lungs at a population level, has to be a massive logistical challenge.
3:08Oh, it's huge. And that scale is really what makes this study statistically powered to actually find these subtle structural variations. I mean, they analyzed RNA sequencing from the whole blood of 3743 subjects in the COPD gene study.
3:24Wow, almost 4000 blood samples. Yeah. And then crucially, they validated and expanded on this using actual lung tissue from 12,241 subjects from the lung tissue research consortium. Over a 1000 lung tissue samples is just phenomenal.
3:39But I want to dig into the computational pipeline here, because I know a lot of our listeners are deep into the bioinformatic side. So we have this mountain of RNA data. How are we connecting it back to those Jewoy zip codes? Because finding alternative splicing events and short read data is, well, it's notoriously difficult.
3:55It is. You're basically relying on tiny reeds that just happened across a splice junction. Right. So how are they quantifying splicing variations accurately enough to run quantitative trait LOSI mapping?
4:07Well, they deployed a really innovative computational tool called leaf cutter? And, uh, the brilliance of leaf cutter is that it completely bypasses the need for a reference transcript on annotation? Wait, really?
4:20It doesn't use a reference? Nope, it doesn't rely on gen code or ensemble to tell it where a squice site should be. Oh, wow. That is critical because, I mean, if you rely on reference annotations, you're basically blind to anything novel or specific to a disease state, right?
4:33Exactly. You only find what you already know exists. So how does Leaf Cutter map it without the reference? It focuses entirely on split reads, so reads that map to 2 distinct locations on the genome, which indicates an excised intron.
4:46Leaf cutter takes all these split reads across the entire cohort and clusters overlapping introns together. Okay. And then it calculates an intron excision ratio. So basically, out of all the transcripts originating from this specific cluster, what percentage use splice junction A versus splice junction B?
5:02Oh, I get it. So it converts structural splicing events into a continuous quantitative variable like ratio that you can then test against genotype data. You got it. And for that testing step, they use Tensor QTL.
5:16Which scales the matrix operations to compute the linear regressions, right? Yeah exactly. It does an incredibly fast on GPUs. But wait, they found over 58,000 splice sites linked to genetic variants in the lung tissue alone.
5:29With a number that massive, how on earth are they controlling for the false discovery rate? Yeah, it's a lot of data Right. I mean, it sounds like it would be incredibly easy to find a statistical ghost in that much multidimensional data, especially when you're comparing 1000000s of SNPs against 1000s of splice clusters.
5:46That is a highly valid concern, and it's exactly why finding SQTL is really only step one. Because, you know, just because a genetic variant drives the splicing change, it doesn't mean that splicing change is actually the cause of COPD.
5:59Right, correlation isn't causation. Exactly. So to prove causality. or at least highly probable shared causality. They used Beesian colloquialization analysis, specifically using the Malike software framework.
6:12Colloquialization is one of those concepts that sounds simple, but is statistically really heavy. If I'm thinking about this, right? It's basically like laying the map of COPD risk directly over the map of splicing errors to find the exact intersections where they match perfectly.
6:29That's a great way to put it. You're layering the maps, but it's not just looking for overlapping P values. You are trying to rule out linkage to equilibrium. Right, LD, can you break that down for us?
6:40Sure. Linkage to equilibrium or LD is the tendency for DNA sequences that are physically close together on a chromosome to be inherited together. Like a package deal. Exactly. So you might have one variant driving this COPD risk and a completely different variant nearby driving the splicing error.
6:58But because they sit right next to each other in the genome, they're in strong LD. They look like the same signal. Ah, so they just happen to travel together on the same chromosomal block, but biologically, they have nothing to do with each other.
7:09Exactly. And that's where Beesian colloquialization comes in. It models different hypotheses and calculates the post-year probability, specifically what is known as post-year probability 4, or PP4. And PP 4 is what exactly?
7:21It's the probability that both traits, So GWA is risk for COPD and the SQTL signal are driven by the exact same shared causal variant. It mathematically separates true pleotropy from mere linkage. That's incredibly elegant.
7:36And to add another layer of stringency, they didn't just rely on the short read leaf cutter data, right? They validated these specific splice junctions using cutting edge Oxford nanopore long read sequencing on the lung tissue.
7:48Yes, which is a massive technical advantage. Standard aluminus sequencing gives you highly accurate, but very short reads. Maybe 150 base pairs. So you have to infer the overall structure of the transcript based on those tiny fragments.
8:00is basically guesswork. Right. But Oxford nanopore literally threads the entire intact RNA or CDNA molecule through a protein pore, measuring the disruption in electrical current. So you captured the full final RNA strands end to end.
8:15You aren't guessing if an Exxon was included. You literally observe the continuous sequence. Exactly. It eliminates so much of the transcript assembly ambiguity that plagues short-rid studies. Okay, so with the maps layered on top of each other and all this validation, what surprising things actually lit up?
8:32Well, the most striking finding wasn't just the sheer number of SkewTLs, but their nature. When they looked at those 58,000 splice sites and lung tissue, roughly 50% of them involved cryptic or completely unannotated splice junctions.
8:48Wait, really? About 50%? Yeah, half of them. Meaning half of the splicing variations happening in our lungs were previously completely hidden from science. That's a mind blowing stat. But I have to push back a little on that.
8:5950% unannnotated. in a well studied tissue like the lung. It sounds high. Is that reflecting actual functional biology or is that just transcriptional noise? We know RNA plummeries isn't perfect. Splicing machinery makes mistakes.
9:12Are we just cataloging the cell's garbage? It is the defining question in transcript comics right now, honestly. Some of it is undeniably just sarcastic noise. But, and what's fascinating here is the colloquialization analysis acts as a biological filter.
9:26When they map these splicing events against the COPD GWS data, they found 38 genomic windows that had a high posterior probability of shared causality. Wow. Yeah, and those 38 windows corresponded to 33 of the original 82 GWA's low sci.
9:42That is a remarkable hit rate. You are essentially taking 33 genomic low sci that we're total black boxes and illuminating the specific transcriptional alterations happening inside them. Yeah. And many of those functional alterations involve these unannotated cryptic sites.
9:56They did. Let's look at one of the most compelling examples from the paper to really illustrate this. The gene FBXO 38. Oh yes. Here's where it gets really interesting. Walk us through the mechanics of what the risk variant is actually doing to FBXO 38.
10:07So the genetic variant associated with COPD risk in this locus directly causes the inclusion of a cryptic exxon in the FPXO 38 pre-MRNA. And this Exxon is not in the standard reference databases at all.
10:20It's a totally rogue piece of RNA code. Exactly. And when this rogue piece is spliced into the mature transcript, it introduces a frame shift that results in a premature stop code on. And a premature stop code on fundamentally disrupts translation.
10:35The ribzone starts assembling the protein, hits this cryptic sequence, reads a termination signal way too early and just aborts. Right, but honestly, it rarely even gets to the translation stage. Eukaryotic cells have a highly conserved surveillance pathway called nonsense mediated MRNADK or MMD.
10:52Ah, the cells quality control system. Exactly. When the cell detects an Exxon Junction complex downstream of a stop codon, it flags that transcript as aberrant. The MD machinery binds to it and rapidly trashes the entire RNA molecule before it can be translated into a truncated, potentially toxic protein.
11:10So the mechanism of the genetic risk isn't just that it makes an altered protein, it's a complete loss of the transcript. The cell just shreds it. The risk variant forces the cryptic exon inclusion, MMD destroys the RNA, and you get a massive depletion of the FBXO 38 protein in the lung tissue.
11:25Yes. Which implies something very important about the normal biology here. If the genetic variant driving COPD risk results in the loss of FBX 038, then the normal full-length FBXO38 protein must be protective.
11:41It's normally protective against COPD. Exactly. It's maintaining long homeostasis. That is incredible. I mean, we know FBX 038 functions as a ubiquitan Legus. It regulates PD1 expression on T cells, which ties deeply into immune regulation and inflammation which are core drivers of COPD.
11:59So by mapping this QTL, we've moved from a vague zip code to a precise mechanism. perfectly illustrates the power of this methodology. And the mechanisms they found are highly diverse. If we look at another major target, they validated beta cellulane or BTC.
12:14The functional consequence is entirely structural, rather than degradative. Right. So BTC doesn't trigger that nonsense mediated decay. The mechanism here involves Exxon 4. What is the COPD risk variant doing to the splicing of beta cellulin?
12:27In the case of BTC, the risk alleal promotes the increased inclusion of Exxon 4. Now, BTC is a critical member of the epidermal growth factor, or EGF family. It's a trans membrane precursor that gets cleaved to release a soluble ligand.
12:40And that ligand that then binds to epidermal growth factor receptors on the surface of epithelial cells in the lung. And that pathway is central to epithelial repair and proliferation, right? Exactly. And Exxon 4 is not just a generic piece of the protein.
12:55It specifically encodes a juxtamembrane stalk, region like a physical loop structure, that anchors the EGF like domain to the cell membrane before it gets cleaved. Oh, I see. So if you force the inclusion of Xon 4, you are lengthening that stock, you are fundamentally changing the three-dimensional structure of the precursor protein.
13:15Yes, you alter the topology. And structural biology dictates function. By changing the length and flexibility of that stock region. You likely alter how efficiently the protein is cleaved or how the resulting laggin binds to the receptors in the lung.
13:27And previous literature already demonstrates that BTC expression drives mucousell metaplasia in COPD patients. So the genetics are driving a structural change in a growth factor, legend, that is directly responsible for the thickened airways and mucous overproduction we see clinically.
13:42It connects the genotype directly to the cellular phenotype. It really does. So what does this all mean for the listener? We now know how these variants break the lung? What do we do with this information?
13:54The broader context here is that approves SKTLs, these splicing errors, are a massive functional driver of complex traits. Steady state EQTLs are just not giving you the full picture? If you're only looking at overall gene expression, you will miss the therapeutic targets.
14:10Take FBXO 38. We now know the protective transcript is being destroyed by nonsense mediated decay. That opens up entirely novel therapeutic factors. What if we design anti-sens oliga nucleotides ASOs that physically block that cryptic splice site?
14:24Oh, wow. So you use an ASO to essentially block the splicing machinery from accessing that rogue cryptic Exxon, the RNA gets spliced normally, you bypass the premature stop code on, and you rescue the production of the protective protein.
14:40Precisely. You are correcting the transcript at the source. It shifts the paradigm from treating the downstream inflammation to repairing the upstream genetic editing error. That is incredibly promising, but I think we also need to address the structural limitations of the study design, right?
14:55Because they utilized bulk RNA sequencing for both the lung and blood cohorts. Yes they did. And we know the lung is a highly heterogeneous tissue. You've got type one and type 2 alveolar cells, ciliated epithelial cells, macrophages, fiber blasts.
15:09By using bulk tissue, isn't that like analyzing a fruit smoothie? That's a perfect analogy. You blend it all up and you capture the aggregate expression. Right. You don't know exactly which individual cell types, the strawberries or the bananas, are actually driving these splicing changes.
15:22Exactly. If a splicing event is uniquely driven by a risk variant in only one rare cell type, say, just interstitial macrophages, that signal might be entirely deleted by the 1000000s of surrounding epithelial cells.
15:36So the colloquializations we see here are likely just the most robust, ubiquitous signals that manage to survive that averaging effect. Which means the true landscape of disease driving SkewTLs is probably much, much larger than what this paper even reports.
15:51Undoubtedly. The obvious next frontier is deploying single cell or single nucleus RNA sequencing to clear up that smoothie problem. We need to isolate the exact cellular niche where these cryptic splicing events are happening.
16:03Yeah, single cell is definitely the future there. There's also a technical confounding factor between the 2 cohorts they use that I noticed. The blood samples utilized total RNA sequencing with ribosomal RNA depletion, but the lung tissue samples utilized polyA selection.
16:18Right, which only captures mature poly-annnihilated Messenger RNA. That is a critical distinction. Because total RNA sequencing captures a much broader spectrum of the transcriptome, right? Including pre-MRNA that hasn't even finished splicing yet.
16:31Exactly. While PolyA selection bias is heavily towards fully processed, mature transcripts. So because they use different RNA prep methods, direct comparisons of splicing efficiency between the blood and the lung data sets have an inherent technical bias.
16:47You might see a splicing event in the blood that seems absent in the lung simply because the polyace selection process filtered it out. Right. It demands a lot of caution. However, I will say, the fact that they still found 38 highly confident collocalized regions, despite these technical and cellular heterogeneities, it really speaks to the immense biological strength of the underlying signals.
17:07It's a profound shift in how we interpret genetic risk. So to wrap this all up, how would you synthesize the core insight from this deep dive in just a couple of sentences? I'd say the central insight is that genetic risk for complex diseases like COPD isn't just a matter of genes being turned on or off.
17:24Often it is driven by subtle alterations in how the RNA is spliced, which can fundamentally change or destroy the resulting protective proteins. That is so well put. And it leaves you wondering, what does this mean for the dozens of other complex chronic diseases where genetics have given us a location but not a mechanism?
17:42Could hidden cryptic splicing be the missing key to unlocking treatments for Alzheimer's, or autoomune disorders, or cardiovascular disease as well? It entirely possible. We're just scratching the surface.
17:53This episode was based on an open access article under the CCBUI 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
18:08If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
18:17Thanks for listening and join us next time as we explore more science, base by base.