This episode examines a large-scale analysis of germline ribosomal DNA (rDNA) variation in ~500,000 UK Biobank genomes that identifies high-confidence rDNA SNVs and indels associating with human complex traits, notably a cluster in the 28S expansion segment ES15L linked to body-size measures.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Yeah, I'm really excited to get into the data for this deep dive.
0:11It is, um, it's a completely fascinating shift in how we think about our DNA. It really is. You know, when we picture the inner workings of ourselves, there is one piece of machinery that always takes center stage, right?
0:25The ribosome. Oh, absolutely. The classic biological assembly line. Exactly. It reads the genetic code and manufactures the proteins that physically build you and um, we usually view these cellular factories as perfectly standardized, identical machines rolling off that assembly line.
0:41Right, just totally uniform. But what if the blueprints for the factory itself have tiny hidden variations that scientists have been completely ignoring for decades? What really happens when the very machines building your biology are slightly customized from person to person?
0:55How could this change your height, your weight, or even your cholesterol? Today we celebrate the work of Francisco Rodriguez Algara, Vardman K. Riken, and their collaborators at Queen May University of London, alongside teams at King's College, London, Oxford, and the NIH, who have advanced our understanding of germ line sequence variation within ribosomal DNA.
1:16And to really grasp the gravity of this shift in how we view translation, we kind of need to unpack the structure of the ribosome itself. Yeah, we learned early in our training that the mature human ribosome is, well, it's a massive ribonuclear protein complex.
1:31It's made of roughly 80 different proteins and 4 distinct ribosomal RNAs. That's the 5S, 18S, 5.8S and 28S subunits, right? Exactly. And because your cells are under this constant pressure to translate proteins, a single gene copy for those RNAs simply cannot meet the transcriptional demand.
1:49It would just be too slow. Way too slow. So the genome solves this bottleneck with tandem repeats. A typical human deployed genome contains anywhere from 200 to 600 copies of this ribosomal DNA or R DNA.
2:00Wow, up to 600 copies. Yeah, and they're organizing these massive clusters across the acrocentric chromosomes. Right, the nuclear organizer regions. But, um, this is exactly where the field has traditionally hit a major methodological wall, isn't it?
2:15Oh, a massive wall. I mean, if you're listening to this and you have ever tried to assemble a genome or map reads from short reads sequencing data, you know that mapive tandem repeats are just an absolute alignment nightmare.
2:30They are notoriously difficult. Standard computational pipelines typically just collapse these arrays in the reference geno. Or researchers just mask them out entirely, right? Because the mapping quality scores just plummet.
2:42Exactly. And the older tech, like commercial micro arrays. They inherently fail to capture them because the probe hybridization cannot distinguish between identical or, you know, near identical repeat copies.
2:53So the prevailing assumption in the field was just that these 100s of copies were essentially homogenized through concerted evolution. Right. We assumed they were uniform. But biological reality is rarely that uniform, is it?
3:03No, it's really not. Due to partial sequence homogenization. We actually have single nucleotide variants and short insertions or deletions scattered across all these 100s of copies in the genome. This raises an important question.
3:15Actually, it raises the central culture of this whole study. Does naturally occurring, inherited germ line genetic variation within human ribosomal DNA actually impact human phenotypes? Like, if the translational machinery itself varies at the sequence level, does the macro level organism actually change?
3:33Exactly. Okay, let's untack this. Imagine you are running a giant restaurant. And in your kitchen, you have 400 copies of a master recipe book. like this analogy. Right. So for decades, the scientists observing the kitchen assumed every single book was perfectly identical.
3:50If they saw a page with a typo, they discarded it, assuming it was just a printing error. Just a glitch. Exactly. But now we are finally reading those typos to see if they actually changed the fundamental flavor of the meal being prepared.
4:03That analogy tracks perfectly with the data challenge this research team had to solve, because to read those sequences accurately at scale, the team leveraged whole genome sequencing data from nearly 500,000 UK biobank participants.
4:18half a 1000000 genomes. That is a staggering amount of data to sift through. It's massive. especially for variants in an unmapped, highly repetitive region. It requires an incredibly robust filter for sequencing noise.
4:30How did they separate biological reality from alignment artifacts? They engineered a highly stringent control by utilizing 49 pairs of monozygotic white British twins. Oh, identical twins. That's clever It is.
4:44And this introduces a crucial metric for this type of research introgenomic variant frequency, or IGF. Right, because we are dealing with multicopy arrays, a variant isn't just a simple heterozygous or homozygous binary, it's a continuous spectrum, you know?
4:58Exactly. If you have 400 copies of the RDNA gene, a specific variant might only appear in, say, 25% of them. And if you are comparing identical twins, their IGF for any true, inherited biological variant should be highly correlated.
5:13Spot on. Twin A and Twin B should both show that specific variant at roughly 25% frequency across their erase. Wait, if these sequences are so repetitive, how do we know a variation is real and not just a glitch in the sequencing machine?
5:28That is exactly what the twin validation caught. They found that about 17% of the identified variants showed extremely poor correlation between twin pairs. So they were false positives. Yes. And when they isolated those discordant false variants, a distinct signature emerged.
5:44They were primarily T to G and A to C transversions. Ah, so the sequencer chemistry itself was creating a mirage. Right. The aluminous short red platforms have known biases in these exact types of genomic environments.
5:57Because it's a 2 channel sequencing chemistry, right? When you have regions with extreme GC content, which ribosomal DNA certainly is, the fluorophore signals can cluster or bleed into each other during the imaging steps on those specific pattern flow cells.
6:11Exactly. The base caller just misinterprets the signal cross talk. That is the exact mechanical failure happening at the sequencer level. Wow. The machine chemistry literally writes in false mutations due to the dense GC repeats.
6:23So by identifying these artifact transversions through the twin discordants, the team established this highly stringent filtering pipeline. They eliminated the machine made glitches and isolated a trustworthy goal standard set of 378 true RDNA sequence variants.
6:40So the monozygotic twins act as an essential filter, leaving you with 378 verified mutations in the ribosomes blueprint. But, you know, a mutation only matters if it actually alters the organism. How did they determine if these specific typos were shifting human biology across the wider biobank population?
6:57Well, they took that filtered list of 378 variants and analyzed their intergenomic variant frequencies across 297,010 unrelated individuals. That's a huge sample size. It really is. And they ran robust linear regressions against a wide array of human complex trades.
7:14Here's where it gets really interesting. Out of those extensive regressions, 34 associations were highly significant. I mean, passing a false discovery rate of less than .01. And those 34 associations map back to just 17 distinct variants.
7:29That is such a tight cluster. Yeah, and the spatial distribution of those variants was completely non-random. All of the highly significant variants were localized to a single subunit, the 28s rabizomal RNA.
7:42So one specific piece of the factory. Exactly. More specifically, they clustered in a hyperactive hotspot around sequence position 10100. And the phenotypes tied to this cluster on the 28S subunit are just striking.
7:55The data showed these variants strongly associate with core body size measurements. We're talking standing height, weight, waste circumference, and even birth weight. Furthermore, they associate with metabolic markers like total cholesterol, HDL, and apolapo protein A.
8:09To understand the magnitude of this effect, we could just look at the height data. The difference in standing height between the highest and lowest deciles of these specific RDNA variant frequencies is about 3 to 4 millimeters.
8:19Three to 4 millimeters from a single locus. I mean, in the context of polygenet traits like height, where 1000s of individual S&Ps contribute tiny fractions of a millimeter, an effect size of that magnitude is wild.
8:33It completely rivals the most strongly associated traditional genetic variants found anywhere else in the single copy genome. It forces a paradigm shift in how we approach missing heritability, doesn't it?
8:44It really does. We have spent years fine mapping the non-repetitive genome to explain complex traits, yet here's a factor of massive effect size hidden directly inside the translational machinery itself.
8:55Let's dig into the physical reality of position 10100 on the 28S subunit. What structural domain does the sequence actually code for on the folded ribosome? That location corresponds to expansion segment 15L, or ES 15L.
9:10Expansion segments, right. Because while the catalytic core of the ribosome is highly conserved across all domains of life, eukaryotic ribosomes have these extra structural regions. Yeah, these expansion segments literally protrude out from the surface of the mature complex.
9:23And ES 15 L in particular has a fascinating evolutionary trajectory. It is known to be greatly expanded in mammals compared to lower eukaryotes. And it is especially pronounced in hominids. The researchers actually pursued that evolutionary angle directly.
9:38Oh did they? Yeah, they compared complete genomes from other apes, including gimpanzees, bonobos, gorillas, orangutans, and Siamangs, just to see if these specific sequences in ES 15 L were shared. Because if ES 15L diverge significantly in our lineage, looking at this variation could tell us whether this is an ancient shared primate feature or a recent evolutionary tweak unique to humans.
10:00Exactly. And the analysis showed that human ES 15 L sequences form a clearly separate distinct cluster from the other primates. Wow, so it's unique to us. Yes. This variation we were seeing, this cluster of variants tied to height and metabolism is entirely species specific.
10:16Humans evolved a unique set of sequence variations right on the surface of this ribosomal expansion segment. So what does this all mean? We have these species specific physical protrusions on our riposomes.
10:28Are these mutated factories actually being used or just sitting idle? That is the definitive regulatory hurdle, right? The researchers had to prove these variants are actually functional. Right, because if a variant copy of the 28S gene is transcribed, The nucleolis might just recognize it as defective and degrade it during rivism biogenesis.
10:46Exactly. So to prove they are used, they generated polysome sequencing data from lymphoblastoid cell lines. Oh, policy and profiling is such an elegant technique. It really is. By using sucrose gradients and tricky gation.
11:00You can physically separate the resting rhabosomal subunits and monosomes from the heavy polysomes. And those polysomes are the dense complexes where multiple ribosomes are actively engaged with a single messenger RNA transcript, right?
11:12Like they're translating it into protein. Exactly. You are literally isolating the active factories. And by sequencing the RNA specifically from that heavy active polysome fraction, they proved that these ES 15L variants are actively transcribed, processed, and physically incorporated into translating rhizosomes.
11:31They aren't silenced at all. They are on the factory floor actively assembling proteins. Okay, but how does a structural tweak on a surface protrusion like ES 15L actually alter systemic traits like height and cholesterol?
11:43Well, if we connect this to the bigger picture. It fundamentally comes down to RNA folding and topographical surface interactions. Right, the 3D shape. Exactly. The researchers utilize 2D structural modeling of the RNA, the modeling demonstrated that different single nucleotide variant combinations within ES 15L actually alter the secondary structure of the segment.
12:06It changes the folding motifs, altering the specific stems and loops of the RNA. Yeah exactly. So it is like, um, swapping out a highly specific nozzle or sensor on a 3D printer. The catalytic core, the machine itself works exactly the same.
12:19But by changing the physical shape of the protrusion on the ribosome's outer surface, you alter the topographical interface. That is the prevailing mechanistic hypothesis. That altered surface protrusion likely changes the ribosome's affinity for specific messenger RNAs.
12:34Oh, so it selectively translates certain MRNAs better. Right. Alternatively, it could modify how the ribosome recruits translation initiation factors, chaperones, or regulatory RNA binding proteins. I see.
12:47So if a specific structural variant of ES 15L increases the translation efficiency of MRNAs related to, say, skeletal growth pathways or lipid metabolism, you directly tune the systemic phenotype of the organism.
13:01Exactly. It firmly introduces the concept of specialized ribosomes to human genetics. The translational pool is not homogeneous. That is amazing. However, we should also clarify the relationship between these sequence variants and the total number of factories.
13:15We established earlier that a genome might have anywhere from 200, 600 R DNA copies. Right. The sheer volume, the overall copy number must also exert phenotypic pressure, right? It absolutely does. Total RGNA copy number is known to associate with various traits, including height.
13:31But are they related to these sequence variants? Well, the statistical models in this study explicitly demonstrate that the sequence variants in ES 15 L and the total copy number operate completely independently of one another.
13:41Wow, they act as independent regulatory dials. Exactly. So it's structural mutation in ES 15L exerts its effect on your height, regardless of whether your overall ribosomal factory count is sitting at the high or low end of the population spectrum.
13:55Correct. The sequence variation provides a qualitative change in ribosomal function, while copy number dictates the quantitative baseline. That is fascinating. It is, but mapping this qualitative landscape is still in its infancy.
14:08And, you know, this study does have distinct limitations we must acknowledge. Right. The twin cohort is the most obvious limitation. Relying on 49 monozygotic twin pairs was mathematically necessary for filtering those sequencer artifacts, but it severely restricts the variant discovery pool.
14:25Entirely, because they needed the twin validation to confidently separate biological signal from sequencing noise, they could only reliably capture the most common variants present in that specific, mostly white British demographic.
14:39Meaning the 378 variants they analyzed are likely just the tip of the iceberg. Oh, absolutely. There is almost certainly a vast landscape of rarer RDNA variance or the population specific variants across different global ancestries that remain entirely unmapped.
14:56Unmapped, but now we know they exist, and we know how to effectively hunt for them. Yes. Moving forward, the field will need to leverage long read sequencing technologies to natively map these arrays without the short read alignment artifacts.
15:10Oh, that makes sense. Furthermore, single cell transcript atomics will be necessary to determine if variant ribosomes are expressed ubiquitously across all tissues, or if different tissues selectively express different ribosome populations to fine tune local translation.
15:25And to definitively prove the mechanics, will need advanced genetic engineering. The ability to introduce targeted edits into mammalian RDNA arrays and directly measure the downstream changes in translation, efficiency, and protein abundance.
15:38That is the experimental horizon. Once we can systematically manipulate these expansion segments in cell lines or animal models, we can map exactly which messenger RNAs are being preferentially translated by which specific ribosome variants.
15:51It totally reframes our understanding of gene expression. Let's synthesize exactly what this research has uncovered. Germline sequence variation in human ribosomal DNA, a massive genetic region historically ignored by standard sequencing pipelines, strongly associates with human complex traits independently of copy number.
16:10Specifically, human specific variants within the 28S expansion segment 15L alter the physical structure of actively translating ribosomes, functioning as genetic dials that tune physical traits like height and weight.
16:23It is a remarkable leap forward for genomics. What does this mean for our understanding of human evolution if the very machines building our biology are uniquely customized to shape who we are? That is exactly the question we need to be asking next.
16:35This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
16:53Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about. Thanks for listening and join us next time as we explore more science base by base.