Law and Burns use a comparative genomics pipeline to date LINE-1 insertions and identify hundreds of chimeric events in mammalian genomes, revealing novel RNA partners and mechanisms that generate composite retrotransposons.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So think about this. 30% of the human Getome, which is around 1000000000 base pairs, was essentially written by a single self-propagating copy paste machine.
0:18Yeah. 1000000000 base pairs, just relentlessly amplifying itself throughout our entire evolutionary history. Right. But what really happens when this ancient genetic machine goes rogue accidentally copying pieces of different genetic constructions and fusing them together into completely new hybrid genes?
0:37Well, I mean, it forces us to look at those 1000000000 base pairs, not just as some inert historical archive, you know, but really as a highly dynamic engine of recombination. Exactly. Today's deep dive is an exploration into a new paper that maps these ancient genetic mashups.
0:52Our mission is to understand the mechanics of how these rare chimeric fusions happen and how they might be the secret engine driving mammalian evolution. It's a massive topic. And before we get into the molecular mechanics, we definitely need to acknowledge the researchers behind this.
1:06For sure. Today we celebrate the work of Chukting Law and Kathleen H. Burns, from the Dana Harbor Cancer Institute and Harvard Medical School, who have advanced our understanding of retro transposans and genomic evolution.
1:19So to start, we really need to look closely at the core machinery driving all this, which is line one. Right, line one, or long interspersed element one. It's a retro transposant that has been active in mammalian genomes for, gosh, over 170000000 years.
1:35Which is an incredibly long time to be jumping around. Oh, absolutely. And to understand how it creates these fusions, we have to look at its primary copying mechanism. So line one produces an RNA intermediate, right?
1:46And that translates into 2 proteins, ORF1P and ORF2P. And ORF 2P is the really critical player here, right? Because it has both endonucleus and reverse transcript taste activity. Exactly. It's the workhorse.
1:58The whole process is called Target Primed reverse transcription or TPRT. Let's actually look at the physical logistics of how that operates on the chrome. So ORF2P has to find a specific sequence, which is usually a T rich region, and physically nick the bottom strand of the DNA.
2:13Yeah, that Nick is the defining event. By cutting the DNA backbone, the endonucleus exposes a free 3 prime hydroxyl group. And then the line one RNA binds to the P rich sequence right at the neck, and the reverse transcript tase uses that free hydroxyl group as basically a primer to start synthesizing a new DNA strand directly from the RNA template.
2:38So it is literally building a new piece of DNA base by base right into the chromosome. Way at it. exactly what's happening. But usually, I mean, it just copies its own line one RNA. The focus of today's deep dive is what happens when something else gets caught in that polymerization process, right?
2:53Creating a bipartite insertion or a chimera. Yeah, and the challenge for genomis hasn't really been conceptualizing that this happens. It's actually finding these chimeras in the genome. Because of how degraded they get over time.
3:05Exactly. 1000000s of years of sequence degradation, newer overlapping mutations and, you know, nested transposable elements, jumping into older ones. It just makes the historical record incredibly fragmented.
3:16So how did researchers normally look for them before this paper? Historically, they relied on a pretty biased approach. using bait sequences. Okay, meaning they had to know what they were looking for. Right.
3:28Like if they suspected a highly transcribed non-coding RNA, like the U6 small nuclear RNA, was forming chimeras, they would computationally scan the genome, specifically for U6 sequences, sitting right next to line one elements.
3:43I mean, the inherent limitation there is obvious, right? You only find the fusions you already suspect exist. If you use U6 as bait, you only catch you six. Exactly. It's like fishing with a lure that only attracts one type of fish.
3:55To map the true landscape of these fusions, the researchers had to abandon the bait sequence method entirely. And that's where their new computational pipeline comes in. Timestamp. Yes, timestamp. Instead of searching for specific known sequences, they basically anchored their search in evolutionary time.
4:11Which is a huge paradigm shift. How does time stamp actually do that? Well, it leverages multiple sequence alignments or MSA. They analyzed 464 genome assemblies across 438 distinct mammalian species. Wow, that is a massive data set.
4:26It is. And by aligning these highly conserved synony blocks across all those genomes, the pipeline identifies orthologous lowci, basically. Regions of the genome that descended from the exact same ancestral sequence.
4:40Okay, so they're looking for insertions that appear contemporaneously, right? Like at the exact same evolutionary time in the specific branch of the evolutionary tree. Exactly. So it's like comparing different editions of a textbook published over 1000000s of years.
4:53Instead of searching for specific typos, you're looking for the exact edition, where a new paragraph and an illustration were suddenly added side by side. That is a perfect analogy, yes. And it allowed the team to pinpoint 12 distinct evolutionary time points for these insertions, completely independent of prior bait sequences.
5:12Wait, I want to challenge the methodology there for a second. When you're looking at an MSA of 100s of species, How does the algorithm definitively distinguish between a chimeric sequence that was inserted at a specific evolutionary node versus, say, a sequence that is much older, but just happen to get deleted in subsequent lineages?
5:30Yeah, that is a really crucial distinction. Yeah. The algorithm relies on the principle of maximum parsimony combined with structural signatures. Maximum parsimony, meaning the simplest explanation is usually the right one.
5:41Right. So if a sequence is absent in distant outgroups, let's say, marsupials and monotremes, but it's uniformly present in a specific clade of placental mammals. It requires way fewer evolutionary assumptions to classify it as a single insertion event at the base of that clade rather than multiple independent deletion events across all the outgroups.
6:02That's a lot of sense. Plus, true line one insertions leave structural hallmarks, like target site duplications, which really help verify that it was a retro transposition event, and not just a complex deletion.
6:14Okay, so taking the blinders off with this timestamp method allowed them to compile this massive catalog of over 700 chimeric events. Yeah, and while they recovered the known ones, like the U6 and 5 S RNA fusions, they discovered a vast array of completely novel line one partners.
6:30Like what? What else did they find? The pipeline identify TRNAs, 28 SRNAs, 7 SLRNAs, YRNAs, and 7 SK RNAs. That's a lot of different RNAs getting hijacked. But the structural data regarding the TRNA fusions is particularly striking to me.
6:45They identify 21 specific TRNA and line one chimeras, but the TRAs Incorporated weren't full-length molecules, they were exactly half-length. Yes, exactly half. Wait, why specifically half a TRNA. I mean, is the cell intentionally cutting them up before they get pasted?
7:01Actually, yes. It connects directly to known biology. Cellular enzymes, like Rnace P2 and Angugenin. They don't just degrade RNA randomly. During cellular stress. They specifically cleave mature TRAs right at the anticodon loop.
7:16Splitting them into discrete 5 Prime and 3 Prime halves. Exactly. Like during oxidative stress, amino acid starvation or even a viral infection. But what is the physiological purpose of cleaving TRNAs during a crisis?
7:29Are they intentionally trying to halt translation? Yeah, that's a primary function. These TRNA halves can interact with the translation initiation machinery or even associate with stress granules, effectively repressing global protein synthesis.
7:42Which can serve cellular energy during the stressor. Right. So what this paper reveals is that the line one machinery is just highly opportunistic. It's capturing these abundant stress induced TRNA fragments floating around the cytoplasm and bringing them into the nucleus to be reverse transcribed.
7:57That is wild. Yeah. But the structural orientation of how they are pasted. That's where the mechanics get really highly complex. These TNA halves are overwhelmingly pasted in the anti-sense orientation.
8:10So basically backwards relative to the line one sequence. Yeah, backwards. And to achieve that, the researchers propose this fascinating mechanism called twin priming. Let's walk through the physical logistics of that, because we have the initial endonucleus nick on the bottom strand priming the line one RNA.
8:27So how do 2 simultaneous copying mechanisms meet in the middle without just unraveling the local chromatin structure? So during TTRT, after the bottom strand is nicked and line one synthesis begins, a 2nd nick eventually has to occur on the top strand to resolve the insertion, right?
8:41To finish the job. Well, in twin priming, that top strand Nick serves as a secondary priming site. While the main line one RNA is being reverse transcribed from the bottom neck, a totally separate RNA molecule, in this case, the TRNA half an eels to the top strand neck.
8:57Oh, wow. So the reverse transcript taste complex begins polymerizing the TRNA sequence from the top neck moving in the opposite direction. Exactly. Two distinct polymerizes or maybe the same complex operating bidirectionally, synthesizing towards each other.
9:11And when the 2 replication forks collide, the host repair machinery ligates them together, locking the TRNA half in a backwards orientation. It is a brilliant deduction from the structural evidence, isn't it?
9:22It really is. And moving beyond small non-coding RNAs, the pipeline also detected 452 chimeric insertions involving alu elements, which is another famous jumping gene. Yeah, Alu is a non-autonomous transposable element.
9:36It doesn't have its own machinery, so it relies entirely on the line one ORF 2P protein to jump. Because Alu elements vastly outnumber line one in the human genome, there must be an intense competition for that ORF 2P machinery.
9:49Absolutely. And the timestamp pipeline provided incredible temporal resolution for these fusions, like 77% of these alu fusions mapped to what they designate as time point seven. Which corresponds to the period just before the divergence of new and old world monkeys, right?
10:06roughly 60000000 years ago. Yep, 60000000 years ago. So if we look at the evolutionary context of the Paleocene epoch back then, what drove this massive surge in retro transposition? Well, the kinetics suggests severe environmental or evolutionary bottlenecks.
10:21When we analyze the specific families of Alu and line one involved, the data shows that in nearly 83% of these cases, the exact families forming the chimeras were experiencing a population search. They were actively transcribing and jumping at the exact same evolutionary moment.
10:36Exactly. They were directly competing for the line one pasting machinery. Which makes sense. I mean, stress induces ritual transposing activity. If early primates were facing massive environmental shifts, their epigenetic suppression mechanisms might have weakened.
10:49Right. And the alu transcripts in line one transcripts were just flooding the nuclear territory simultaneously. Aggressively competing for limited ORF 2P binding sites. And when both transcripts interact with the same TPRT complex at a genomic Nick, you get these massive fusion events.
11:04Yeah, the density of transcripts directly influences the error rate of the reverse transcription process. But here's the thing. Because timestamp was unbiased, it didn't just find repetitive junk DNA fusions.
11:16Right. It actually found 17 discrete instances where actual protein coding MRNAs or long non-coding RNAs fused with line one. Which is huge. The specific examples they mentioned, the MAP 3K13 and XHIT genes are remarkable because of precisely which parts of the host genes were captured.
11:34In these instances, the very 1st exon of the host gene was spliced flawlessly onto a nearly full-length protein coating line one element. Yeah, we are looking at a mechanism called transsplicing there.
11:44But how does the host spicy some, which is usually so highly regulated, confuse a line one transcript with the host gene's own downstream exxons? Well, splicing is essentially driven by physical proximity and sequence motifs.
11:56Mostly spliced donor and acceptor sites. In the densely packed environment of a transcription factory inside the nucleus, multiple genes and transposable elements can be transcribed in really close spatial proximity.
12:09Okay, so they're literally just bumping into each other. Pretty much. And if a line one transcript possesses a strong splice acceptor site, and it is floating near an actively transcribing host gene, like MAP 3K13, the splice system complex can just mistakenly ligate the 5 prime splice donor of the host's 1st Exxon directly to the line one RNA.
12:30So it's basically like a spam email accidentally acquiring the official letterhead of the CEO. Because it has the CEO's header, it completely bypasses the spam filters. I love that. Yes, exactly. By stealing the 5 prime end of a gene.
12:42The line one sequence is essentially co-opting the host genes promoter and regulatory binding sites. It acquires this highly complex, preapproved structural format that dictates exactly when and where the transcript should be expressed.
12:54And that promoter acquisition alters the regulatory landscape completely. By capturing it, the line one element frees itself from its own internal regulatory constraints, and it falls under the control of the host's regulatory network.
13:07Which brings us to the RAP1 GDS1 case, which I think serves as a really powerful demonstration of promoter co-option in action, like a zombie line one. Oh, totally a zombie line one. So the researchers found a specific line one element, an L1PA2 family member, situated right inside an intron of the RAP1 GDS1 gene.
13:27And this particular L1PA 2 had a massive deletion at its 5 prime end, right? entirely removing its internal promoter. Right. So on its own, this line one sequence was transcriptionally dead, completely dead.
13:39It lacked the capacity to recruit RNA polymerase and initiate transcription at all. Get it hijacked the highly active promoter of the RAP1 GDS1 host gene, because it resides within that gene, the powerful promoter drives expression right through the intron.
13:53Exactly. And through splicing, that expression is just force fed into the truncated line one. The fusion transcript restores the line one's ability to produce functional proteins, essentially reviving a dead sequence.
14:04And here's the crazy part. RAP1GDS1 exhibits high expression levels in specific tissues, predominantly the brain and the testis. So the host genes promoter profile restricts this revived line one activity precisely to those tissues.
14:18Yes. And the researchers even validated this by querying modern long read sequencing data sets. They found definitive evidence of this exact RAP1 GDS1 line one chimeric RNA actively circulating in human brain tissue today.
14:32That is just incredible to think about. We really have to analyze why this matters. Like why promoter swapping is such a critical survival mechanism for retro-transposans, especially considering the intense defenses our host cells deploy against them.
14:44Yeah, the host cell utilizes very sophisticated epigenetic silencing machinery to suppress line one activity. Complexes like the hush or human silencing hub, alongside targeted DNA methylation, they're specifically designed to recognize the structural signatures of the line one promoter.
15:00Basically locking it down. Exactly. Methyl transfer raises add methyl groups to the cytocenes in the line one promoter, which recruits reader proteins that tightly coil the local chromatin into heterochromatin.
15:12This physically compacts the DNA, totally blocking RNA polymerase from accessing the sequence. So the silencing machinery acts as a structural padlock, but the target of that padlock is the native line one promoter.
15:25If a line one element utilizes transplicing or intronic trapping to swap its native promoter for a diverse RNA or a host genes promoter, it bypasses the padlock entirely. Exactly the case. The host cell cannot simply methylate and silence the RAP1 GDS1 promoter because the cell absolutely requires the RAP1 GDS1 protein for neural development and function.
15:47So by hiding behind an essential host promoter, the line one element evades heterochromatin formation and maintains its expression, it leverages the cell's own essential architecture against the silencing mechanisms.
15:58It's remarkably insidious. However, I should point out, the researchers are careful to note the limitations of the current study. Right, because the timestamp pipeline relies heavily on the quality of the underlying genome alignments.
16:08Yeah, when you are constructing alignments across highly divergent clades, like comparing placental mammals to marsupials, homologous repetitive regions can be misaligned, creating computational artifacts.
16:20So what's the next step? Well, the bioinformatic evidence for transsplicing and twin priming is really robust, but the next vital phase requires experimental validation, recreating these specific cellular stressors and controlled in vitro environments to directly observe the kinetics of twin priming or the specific error rates of the splices some during transsplicing that will be critical to verifying these mechanisms.
16:44To really distill the central insight from all of this. Line one is not merely a mindless repetitive element passively accumulating mutations. It is a highly dynamic engine of recombination. By fusing with diverse host RNAs, from stress cleave TRNA fragments to complex protein coating MRNAs, line one, constantly generates novel genetic architectures that have profoundly shaped mammalian evolution.
17:07Absolutely. It totally challenges the paradigm of genomic stability, demonstrating that the genome is continually experimenting with its own structural components over 1000000s of years. So what does this mean for how we view the so-called junk DNA inside our cells?
17:21If these retro transposition events are highly active in neural tissue driven by promoters like REP1 GDS1, we must consider the ongoing impact. Could these rogue copy paste fusions be happening in your brain right now, subtly altering the genomic architecture of individual neurons are contributing to the unique, complex wiring of your mind?
17:42That is a fascinating thought to leave on. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
17:54If you enjoy this, follow or subscribe in your podcast app and leave a five-star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
18:07Thanks for listening and join us next time as we explore more science based by base.