Across 34 complex traits and disorders, a MiXeR-based framework partitions SNP heritability over 74 functional annotations and finds that exons carry only a minority of it, and steadily less as a trait becomes more polygenic. Exonic heritability falls from about 22 percent in less-polygenic somatic diseases and biomarkers to about 13 percent in highly polygenic psychiatric and cognitive traits, intergenic heritability rises by the same logic, and intronic heritability stays put. A new annotation contribution score shows the same axis in the annotations themselves: highly polygenic traits load on conservation and variant-effect scores, less-polygenic traits on promoter, transcription and chromatin marks.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Base my base is now on YouTube too, at Base by Base, where every episode gets a video with chapters and the full description come subscribe.
0:15So I want you to take 2 identical twins, for example. You know, they share the exact same DNA, right? Right. The exact same genetic words. Yeah, the exact same words written into the billions of letters of their code.
0:27They grew up in the same house, eat the same meals, go to the same schools. With all the exact same environmental baselines, yeah. Exactly. Yet, as they reach early adulthood, one twin develops schizophrenia while the other does not.
0:40Which is just, I mean, it's wild when you really think about it. It is. And for decades, scientists tore apart those genetic words looking for a typo. Like the sequence, the genes, they looked for the smoking gun, you know, the broken gear in the machine that would explain why human biology can diverge so wildly.
0:58Right. They wanted a clear cause and effect. But more often than not, they were looking in the wrong place entirely. When we think about our DNA shaping who we are, we naturally picture those genes, the tiny fraction of our code that actually builds proteins.
1:12Right, the stuff that actively makes things. Yeah. But what if the traits that make us the most profoundly complex, like our cognition or our risk from mental health disorders aren't driven by a few broken genes?
1:24Oh, this is the big question. Right. What if they're driven by thousands of tiny, invisible tweaks hidden in the vast, unlit spaces of our DNA? You know, the regions we once arrogantly dismissed as junk.
1:36The premise of that question just completely changes the entire foundation of modern genetics. We spent, it's better part of a century, basically operating under a factory floor model of biology. A factory floor.
1:49Like how so? Well, you know, a gene makes a protein, the protein does a job. If the job isn't done right, you go find the faulty gene. It's very mechanical. Right, like looking for a broken part on an assembly line.
1:59Exactly. Yeah. But looking at the spaces between those genes forces us into a completely different paradigm, we have to stop looking for a single broken cog. And start looking at what exactly? We have to start trying to comprehend an unimaginably vast, interconnected network of regulatory switches.
2:16Switches that do what? They can turn the volume up or down on those teams, depending on the tissue, the time of day, or, you know, the stage of human development. Wow. How could this change the way we search for cures to psychiatric conditions?
2:30And what really happens when the instructions for a trait are scattered across the genomic wilderness? It's a huge shift in how we understand disease. It really is. Today we celebrate the work of Julian Furr, Alexander Frey, and their colleagues at the University of Oslo and the J. Craig Venter Institute, who have advanced our understanding of the genetic architecture of complex traits.
2:50It's phenomenal piece of research. Truly. Their work was published in the American Journal of Human Genetics on October 1st, 2026. And the deeper we go into this analysis, the more it will completely alter how you, the listener, visualize the blueprint of human life.
3:06Yeah, the magnitude of what Fur, Frey, and their team have achieved here really requires us to understand a historical bias in genetics. Okay, lay it on us. What's the bias? It's what we often call the lamp post problem.
3:21Oh, the lamp post problem. I love this analogy. Right. So picture someone looking for their lost car keys at night under a street lamp. A police officer walks up and asks, is this where you lost them? And the guy says, no, right?
3:33Yeah, the person replies, no, I lost them in the park, but the light is better over here. That is just so perfectly human. It is. And for a very long time, genetic research has been looking for the keys to human disease under the lamp post of the XO.
3:48Okay, let's define the egg zone right away for those who might just be catching up on this field. Good idea. So the XM is simply the collection of all the Exxons in the genome, which are the specific sequences of DNA that code for proteins.
4:00Right, the actual parts that do the heavy lifting. Exactly. They are the actual recipes for the physical building blocks of your body. And because they spell out the recipes for proteins, the light is indeed very bright there, to use the analogy.
4:14Because it's easy to see when something goes wrong. Precisely. If a single letter in an Exxon is mutated. The resulting protein is misshapen. We can track that. We can physically see how that misshapen protein causes sickle cell anemia or cystic fibrosis or Huntington's disease.
4:32So the cause and effect are direct. Yeah, they're observable and mechanistic. It is highly rewarding science because it provides clear, straightforward answers. But the problem with the lamp post is that it ignores the sprawling darkness of the park.
4:45A huge massive park. Right, which, in this analogy, is the other 98% or so of the human genome, the non-coding regions, the so-called junk DNA. The term jink DNA is arguably one of the greatest linguistic missteps in the history of science.
5:01It really is terrible branding. Right. Just because somebody doesn't code for protein does not mean it lacks function. Yeah, that's just arrogant. It was. But beyond that bias, the core issue this paper addresses is a concept called polygenicity.
5:15Oh, polygenicity. This is big. Yes. Polygenicity is basically the measure of how many separate genetic variants. How many individual tweaks in your DNA contribute to a single trait or disease? I think this is a crucial concept to anchor before we go further into this deep dive.
5:32Let's compare something with low polygenicity to something with high polygenicity. Sure, a classic example of extremely low polygenicity or a monogenic trait is something like cystic fibrosis, which we just mentioned.
5:44Because it's just one gene. Right, exactly. It is caused by mutations in a single gene, the CFTR gene. If you have the mutations, you have the disease. It is just a binary switch. Right. But when we look at complex human traits, say your height, your blood pressure, or your risk for schizophrenia, there is no single height gene, is there?
6:03No, not at all. There's no single schizophrenia gene either. These trace are shaped by thousands, sometimes tens of thousands of different genetic variants scattered across your entire genome. So it's a massive group effort.
6:14Yeah, each individual variant might only increase your height by, like, a millimeter or increase your risk for a psychiatric condition of a percent. Yeah. That is high polygenistic. Wait, let's unpack this with another analogy.
6:28If the genome is a massive instruction manual. The exxons, the protein coding genes are the actual words on the page. Right, the text itself. Yeah. And for a long time, we only focused on the words. We thought, you know, if the words are spelled correctly, the manual is perfect.
6:43Which makes sense on the surface. It does. But the rest of the genome, that massive non-coding space is the formatting. It's the bold text, the italics, the cross references, the paragraph spacing, the table of contents.
6:57And we are finally asking the question. How much does the formatting actually dictate the final product, especially when the manual is incredibly complex. That is a phenomenal way to conceptualize the biology, honestly.
7:09The non-coding regions act as the formatting that tells the cell when to read the words, where to read them, and how loudly to pronounce them. Okay, give me an example. Well, if you have the word grow in the manual.
7:20The formatting tells the bone cell to read it during puberty, but tells the brain cell to ignore it entirely. Oh, wow. So the word is the same, but the formatting changes its impact completely. Exactly.
7:31And what the researchers at the University of Oslo have done is tackle one of the most massive, unresolved questions in modern genomics. They wanted to resolve where heritability sits across the genomic landscape, based specifically on a treats polygenicity.
7:47Meaning they wanted to see if a highly complex, highly polygenic trait, like schizophrenia, relies on different parts of the instruction manual than a simpler, less polygenic trait, like, say, cholesterol level.
7:58That's it, exactly. They set out to map the distribution of genetic risk. Does heritability remain static in the protein coating regions, regardless of the trait? Or does it shift into the non-coding wilderness as traits become more complex?
8:11Which brings up a logistical mountain? I mean, how do you map the entire genomic wilderness for multiple complex trait simultaneously? You can't exactly bring 100,000 people into a lab and run a controlled experiment on their brains.
8:24No, you definitely cannot do that. So wait, just to be precise, no one was in a lab drawing blood or running physical experiments on patients for this specific paper, right? Absolutely not. This is entirely a statistical partitioning of heritability.
8:38Okay, so it's all math. Yes, it is a mathematical modeling of data we already have. They use something called GDVUSS summary statistics. G-W-S-E-S, right. Let's slow down and really dissect what a GA is, because we hear that acronym thrown around constantly in science news.
8:56Yeah, it stands for Genome Wide Association study. Right, but the actual mechanics of it are incredibly vital to understanding why this paper is so groundbreaking. G-Waves is essentially a massive statistical census of the genome.
9:09Imagine you want to find the genetic variants associated with bipolar disorder. Okay, so you need a lot of people for that. You do. You gather the DNA of, say, 50,000 people diagnosed with bipolar disorder, and you gather the DNA of 100,000 people who do not have it.
9:23They serve as your control group. Got it. 50,000 with 100,000 without. Right. And then you scan millions of single letter variations in their DNA. These variations are called single nucleotide polymorphisms or S&Ps.
9:36Pronounced snips. Yeah, exactly, snips. So you are basically putting these 2 massive populations side by side and asking a computer program to find any single letter in the genetic code that shows up more frequently in the bipolar group than in the control group.
9:50That is the essence of it, yeah. If a specific T instead of a C at a very specific location in the genome shows up in 40% of the bipolar group, but only 20% of the control group, the GWS flags that variant as being statistically associated with the disorder.
10:05But it doesn't mean that tea causes bipolar disorder all on its own. Right. It doesn't mean that variant causes the disorder by itself, but it is a statistical red flag on the genomic map, indicating that this neighborhood of DNA is somehow involved.
10:19Okay. And when you say summary statistics, you mean they didn't take the raw individual DNA data of all these hundreds of thousands of people? No, no, that would be a privacy and logistical nightmare. Yeah, so they just took the final tallies, the lists of flags and their statistical weights from previous studies.
10:35Exactly right. They gathered these GW summary statistics across 34 different complex traits. 34. That's a lot of data. It's a massive amount of data. And this was a highly deliberate choice to ensure they were capturing a broad spectrum of human biology.
10:52They didn't just look at psychiatric condition. What else did they include? They included anthropometric measures like heightened body mass index. They included physiological biomarkers like cholesterol, triglycerides, and systolic blood pressure.
11:05Right, the basic bodily functions. Yeah. And of course, they included highly complex cognitive and psychiatric traits like schizophrenia, educational attainment, major depressive disorder, and general cognitive ability.
11:17So they have this mountain of statistical flags for 34 different human traits. How do they even begin to organize it? They use a tool called the mixer-based framework. Mixer. Okay, can you walk us through what that actually does in plain English?
11:32Sure. Mixer is an advanced statistical modeling tool. Its primary job is to estimate 2 things. Polygenicity, which we defined earlier as the number of variants influencing a trait, and the effect size of those variants.
11:45Okay, simple enough. But it goes a step further, by accounting for something deeply fundamental in genetics, called linkage disequilibrium. Oh boy. Linkage dis-equilibrium. That is a phrase that terrifies undergraduate biology students.
12:00It does have that reputation. Let's break that down because it's the reason mapping the genome is so incredibly difficult in the 1st place, isn't it? It really is. It sounds intimidating, but the concept is actually quite intuitive once you get past the name. Linkage to equilibrium, or LD is essentially the neighborhood effect of DNA.
12:19The neighborhood effect. Yeah. When you inherit DNA from your parents, you don't inherit a completely shuffle deck of individual letters. Right, it's not totally random. No, you inherit large chunks of DNA intact from your ancestors.
12:32Like getting a whole chapter of the instruction manual pasted in at once rather than individual words. Exactly. Because these chunks are passed down together. The genetic variants within that chunk travel together through generations.
12:44Ah, I see where this is going. So, if a GWU flags a specific variant, as being associated with a disease, the reality is that the variant might just be a harmless neighbor living on the same street, as the actual disease causing variant.
13:00Wow. So because they are inherited together in a block, the GWS flags the whole neighborhood. Right. Linkage to equilibrium is the statistical correlation between these neighbors. The mix of framework is sophisticated enough to account for this neighborhood effect.
13:15Meaning it can filter out the innocent bystanders? Essentially, yes. It allows the researchers to estimate the true number of independent causal variants driving a trait rather than just counting every single correlated flag on the block.
13:27This is why we need supercomputers in advanced mathematics for this stuff. You are trying to find the one house on the block that is causing the problem when the entire block lights up on your scanner.
13:38Exactly. So mixer helps them estimate the true polygenicity. But the overall goal of the paper, to figure out where this heritability sits. Right. Right. They had to map these statistical flags onto actual biological geography.
13:52To do this, they partition the genome into 74 functional annotation. Let's define functional annotation for the audience. Think of the genome as a map of a massive country. A functional annotation is like a zoning map overlaid on top of it.
14:06Okay, like city planning. Yeah, some zones are residential, some are commercial, some are industrial. In the genome. The researchers categorize the DNA based on what role it plays in the cell. What kind of zones are we talking about here?
14:18Well, the most basic structural categories are the ones we've touched on. Exxons, which are the residential zones where the proteins are actually built. Then you have introns, which are the non-coding sequences trapped inside the genes, basically interrupting the Exxons.
14:32Like little empty lots between the houses. Sort of, yeah. And finally, you have the energetic regions, which are the vast seemingly empty highways of DNA completely between the genes. But you said 74 categories?
14:44That means they went way, way deeper than just Exxon's introns in intergenic space. What else were they looking for in this zoning map? Oh, they went incredibly deep. They mapped out the regulatory machinery.
14:56They look at transcription factor binding sites. Okay, transcription factors. Those are the proteins that turn jeans on and off, right? Yes. They are specialized proteins that physically attach to the DNA to turn jeans on or off.
15:08They are the fingers pressing the keys on the genomic piano. The specific sequences where they attach are crucial regulatory zones. That makes total sense. What else? They also looked at evolutionary conservation scores, meaning sequences of DNA that remain identical whether you look at a human, a mouse, or a fish.
15:28Oh, wow. So if a fish and a human have the exact same sequence. It means it's vital. If evolution refuses to change a sequence over hundreds of millions of years, you can bet it has a critical function.
15:38Because if you change a fundamental piece of code and the organism dies, evolution just doesn't pass it on. Exactly. So conserved regions are highly protected zones. You mentioned earlier, they also looked at things like chromatin states and epigenetic markers.
15:52This is where the physical structure of DNA becomes really important, doesn't it? It is paramount. We often visualize DNA as a neat double helix floating freely, but in reality, to fit 2 meters of DNA into a microscopic cell nucleus, it has to be tightly schooled around proteins called histones.
16:12Right, like thread on a spool. Yeah. And this complex of DNA and protein is called chromatin. A chromatin state refers to how tightly or loosely the DNA is spooled. I always think of it like a zipper on a jacket.
16:25If the chromatin is tightly zipped up, the cellular machinery can't get in there to read the genes hidden inside. But if the chromatin is unzipped, the genes are exposed and can be activated. That zipper mechanism is driven by epigenetic markers.
16:37These are tiny chemical tags, like methyl groups or recetal groups, that are attached to the DNA or the histones. But they don't actually change the DNA letters, right? No, they don't change the underlying letters of the genetic code at all.
16:50They just act as the physical tags that tell the cell whether to zip or unzip that section of the chromatin. Wow. So by including these in their 74 functional annotations, the researchers were essentially mapping out not just the words in the manual, but the physical bookmarks and highlights that control how the manual is read.
17:09Exactly. It's an incredibly comprehensive map. Okay, so we have the statistical flags from 34 traits on one hand, and this incredibly detailed zoning map of 74 functional regions on the other. They are laying one map over the other to see where the flags land.
17:23But there is a massive mathematical trap here, right? Because these zones are not all the same size. Oh, it is the most common pitfall in genomics. Yeah. Let's say you overlay the max and find that 2% of the heritability for a trait lands in a specific regulatory zone.
17:39Is that important? I mean, 2% sounds small. Well, it entirely depends on the size of the zone. If that regulatory zone only makes up 0.one% of the entire genome. Then capturing 2% of the heritability is a huge deal.
17:53Ah, I see. It means that tiny zone is highly enriched. It is pulling far above its weight class. Right. It's about population density. If I tell you, I found 10 millionaires in a one square mile area, you'd be impressed if that area was an empty glacier in Alaska.
18:09Right. But you wouldn't care at all if that area was downtown Manhattan. The baseline expectation is totally different. The Manhattan versus Alaska analogy is perfect. If an annotation covers 50% of the genome, but only contains 10% of the heritability, it is actually depleted.
18:24It is a genomical Alaska holding a tiny fraction of the expected value. So it's basically empty space compared to what you'd expect. Exactly. And historically, statistical models have struggled to compare these regions fairly because the tiny, highly enriched regions can artificially skew the results, simply because their signal density is so high.
18:44So how did the researchers fix the Manhattan and Alaska problem? This is the critical innovation of the study. They introduced a metric called the annotation contribution score, or ACS. ACS. What does that do?
18:57Instead of just asking, how much heritability is in this zone, the ACS asks a much more sophisticated question. How much does knowing about this specific zone actually improve our overall statistical model relative to the zone's physical size?
19:12Oh, so it's essentially punishing a region for being huge and rewarding a region for being small to level the playing field? For Christ sakes, it allows you to actually compare an Exxon to a massive energetic desert fairly.
19:25It measures the true information game. That is so clever. It really is. It allows us to objectively rank the importance of these 74 different annotations without size bias muddying the waters. And this brings us to the core findings of the paper.
19:39Yes, let's get into those numbers because once they applied this massive computational framework, the number they produced completely upended the traditional view of human biology. They absolutely did.
19:49Let's start by shining a light on the lamppost itself. The Exxons, the actual protein coding words. What did the data actually say about their role across these 34 traits? We must be very precise with the data here.
20:01The researchers confirm that Exxons make up only 2.55% of the total base pairs in the human genome. So 2.5% of the physical space, tiny. An incredibly tiny fraction of the physical DNA. Yet when they partition the heritability across the 34 trades, they found that these Exxons carry a mean of 14.52% of the total heritability.
20:23Wait, let's just make sure that fully sinks in for you, listening. The protein coating genes. The literal blueprints for the physical machinery of your body, the things we have spent billions of dollars and decades of research obsessing over account for less than 15% of the heritable genetic variation for these complex traits on average.
20:40Exactly, 14.52%. Yeah. It is a sobering statistic for the classical geneticist. I mean, the Exxons are indeed enriched. They carry 14.52% of the heritability despite only taking up 2.55% of the space, which proves they are functionally vital.
20:54Right. Right. They are punching above their weight class. Yes. But as you point out, it leaves roughly 85% of the heritability sitting out in the non-coding dark matter of the genome. And this is where the paper stops being just a statistical census and becomes a profound story about evolution and complexity.
21:10Because that 14.52% is just an average across all 34 traits, right? Right. It just a me. The reality is that the genome handles different types of traits in drastically different ways. This is where the researchers uncovered a definitive, measurable sliding scale based on polygenicity.
21:29They separated the traits into categories based on how complex they were. Okay, let's look at that sliding scale. What on the low complexity end? On one end, you have less polygenic somatic diseases and biomarkers.
21:39These are the mechanical physiological functions of the body, things like circulating sex hormone binding globulin, your systolic blood pressure, or your height. Basically, the nuts and bolts of keeping a human body physically running.
21:52Precisely. For those physiological somatic traits, the exotic care ability average is about 22%. 22%. Okay, that's a significant chunk. It is. The genes themselves are carrying a substantial burden of the variance there.
22:05But as you slide down the scale toward traits that are highly diffused, highly complex, and intensely polygenic. Like what, specifically? Specifically psychiatric disorders like schizophrenia and major depression or cognitive traits like educational attainment, the architecture fractures.
22:23The exotic heritability plummets to about 13%. This was the moment in the paper that genuinely gave me pause. The more complex the human trait. The more it involves the brain, behavior, and a higher order function, the less it relies on the actual protein coating genes themselves.
22:41completely counterintuitive. Right. You would think a more complex machine needs more complex parts. But instead, the genome uses the exact same parts and just vastly complicates the instruction manual on how to assemble them.
22:53They actually quantified this shift, didn't they? They did. They mapped the slope mathematically. They found that for every tenfold increase in polygenicity, meaning every time a trade becomes an order of magnitude more complex in terms of the number of variants involved, the exotic fraction of heritability falls by 4.38 percentage points.
23:11Falls by 4.38 percentage points. So the genes are literally bleeding influence as the trait gets more complicated. That's good way to put it. Where does that influence go? It shifts directly into the energetic spaces.
23:24In perfect symmetry with the Exxons dropping, the energetic fraction, the heritability found in the vast deserts of DNA between the genes. Rises by 4.87 percentage points per tenfold increase in polygenicity.
23:38rises by 4.87. It like a seesaw. It is very much like a seesaw. As traits require more and more genetic inputs. The genome is essentially outsourcing the complexity. It pushes the control mechanisms out of the genes and into the distant regulatory regions.
23:51The intergenic dark matter becomes the master conductor of the symphony. Beautifully said. The energetic regions harbor those transcription factor binding sites, those distal enhancers, they can loop across huge physical distances to turn a gene on or off.
24:06So the action is happening far away from the gene itself. Yes. For complex traits, the variation isn't in a broken protein. The variation is in a switch located a million base pairs away that slightly delays the production of a protein during a critical window of brain development.
24:23Now, here is the mystery that I want to spend some serious time on because it feels like the missing piece of the puzzle. We have the Exxons going down, we have the energenic space going up. But what about the introns?
24:35Oh, the introns. We defined them earlier as the non-coding sequences that are wedged physically inside the genes? They are transcribed by the cell, but they are usually cut out and thrown away before the protein is built.
24:46They are the cutting room floor of the genetic code. Right, they are spliced out. So if the extremes of the genome, the exxons, and the intergenic spaces are totally shape shifting based on complexity.
24:57What happens to the introns? This is perhaps the most fascinating finding in the entire study. The intronic slope is not significant. Let me just clarify that. It doesn't drop like the exons, and it doesn't rise like the energetic space.
25:10It just stays perfectly flat. Statistically, the change is not significant, and the sheer volume of heritability they carry is staggering. How much do they carry? Introns alone carry roughly half of the heritability across these traits.
25:23Yes, roughly 50%. And that massive share stays remarkably stable, regardless of polygenicity. Whether you are looking at a relatively simple blood biomarker or a highly complex psychiatric condition, the intronic regions consistently hold down about half of the genetic influence.
25:40Why? I cannot wrap my head around this. If evolution is aggressively shifting control from the words to the distant formatting, based on how complex a trade is, why is the middle ground, the introns just anchored in place at roughly 50%?
25:54What is so universally essential about the cutting room floor? To answer that, we have to deeply investigate what introns are actually doing. Because they are definitely not just trash. The researchers recognized this mystery and broke the intronic sequence down further.
26:10Okay, how did they break it down? They separated the introns into conserved intronic sequence. Those evolutionary protected sequences we discussed earlier, that are shared across species, and non-conserved intronic sequence.
26:22Ah, looking for the parts that evolution refused to change. Exactly. And this is where the nuance appears. While the overall intronic fraction of roughly half is stable, the heritability within the conserved intronic sequences actually rises as polygenicity increases.
26:38Wait, really? So evolution is aggressively protecting specific chunks of non-coding code inside our genes, and those specific chunks become vastly more important, the more complex the human brain gets.
26:49Yes. What are those conserved sequences doing? They are the master directors of a process called alternative splicing. Okay, we need a full ELI 5. Explain it like I'm 5 for alternative splicing because I suspect this is the key to human complexity.
27:03It absolutely is. Let's consider a humbling biological fact first. Humans have roughly 20,000 protein coating genes. Okay. 20,000 genes. Seems like a lot maybe. Well, a microscopic nematode worm, C elegance, which is about a millimeter long and has exactly 302 neurons, also has roughly 20,000 protein coating genes.
27:24Oh. So gene count alone, absolutely does not equal complexity. Otherwise we would be indistinguishable from a soil worm. Exactly the point. How do humans achieve our staggering cellular diversity, particularly the 86000000000 neurons in our brain with the same number of genes as a tiny worm?
27:41How do we do it? The answer is alternative splicing. When a gene is read by the cell, the entire sequence, exxons and introns together is copied into a temporary molecule called RNA. Okay, so you make a temporary copy of everything.
27:52Right. But the cell's machinery then has to edit this RNA. It literally cuts out the introns and splices the exxons together to make the final protein recipe. This is the wardrobe analogy we've played with before on the show.
28:03The Exxons are the individual pieces of clothing in the wardrobe. A shirt, a jacket, a pair of pants. The RNA transcript is the whole closet. Yes. And alternative splicing is the stylist. Depending on the signals the cell receives, the stylist can choose to include Exon one, X on 2 and X on 3 to make a specific protein.
28:25Like putting together a specific outfit. Right. But in a different cell type or at a different time, the stylist might choose Exxon one, skip Exxon 2 entirely, and splice directly to Exxon 3. Oh wow. By mixing and matching the exxons, a single gene can produce dozens, sometimes hundreds of drastically different protein variants which are called isoforms.
28:46So one gene isn't just one protein. One gene is a modular toolkit that can build multiple different tools depending on how you assemble the parts. Precisely. And what controls the stylists. What tells the machinery which exxons to keep and which to skip?
28:58The instructions for the stylist are hidden within the intro. It all comes back to the intons. It does. Those conserved intronic sequences are the docking stations for the splicing machinery. They dictate the cutting and pasting.
29:11And alternative splicing is utilized more heavily in the human central nervous system than anywhere else in the body. Because the brain is so complex. The brain relies on this mechanism to generate the immense molecular diversity needed to wire billions of neurons together in complex circuits.
29:27This brings everything together beautifully. It makes perfect sense that as a trait becomes more highly polygenic, as it starts to involve the mind boggling complexity of human cognition, psychiatric vulnerability, and brain development.
29:40The heritability would shift toward those conserved intronic regions that regulate alternative splicing. You need the stylist to be working overtime to generate the complexity of the brain. It highlights a profound evolutionary strategy.
29:53You don't need to invent new genes to build a human brain. You just evolve a vastly more complicated regulatory system to splice, delay, and modulate the genes you already have. You reuse a limited tool kit in incredibly diverse cellular context.
30:07Okay, I have to play doubles advocate here for a second. I want to push back on this because I know what someone listening to this might be thinking right now. All right, let's hear. If I'm hearing that 85% of heritability is in the non-coding space and that the Exxons, the genes themselves drop down at just 13% for cognitive and psychiatric traits, the logical leap is to say, genes don't really matter for human psychology.
30:32all just formatting. How do we reconcile this data with the reality of genetic diseases? That is a critical point of friction, and I am so glad you brought it up. We must be incredibly precise here, and the authors of the paper are emphatic about this limitation and interpretation.
30:48Do not conclude that genes or excellence do not matter. Say that again for the folks in the back. Do not conclude that genes or exxons do not matter. That is a dangerous misinterpretation of the data. So how do we explain it to a patient?
31:01If a patient has a genetic disorder, you can't exactly tell them the genes are just a tiny piece of the pie. You have to separate the concept of rare, severe mutations from common, complex variation. Okay, what's the difference?
31:14If you take a sledgehammer, to the engine block of a car, the car will not run. A severe rare mutation in an Exxon is a sledgehammer. It breaks the fundamental structure of the protein. Like the cystic fibrosis example from earlier.
31:29Exactly. Conditions like early onset Alzheimer's or severe neurodevelop metal disorders are often caused by these rare exotic mutations that have massive, devastating effect sizes. The word in the instruction manual is entirely scrambled, so this sentence makes no sense.
31:46The manual is just broken. Precisely. But this study is looking at highly polygenic traits within the general population. We are talking about common variations. Single letter changes that are present in 10%, 20%, or 30% of the population.
31:59Okay, so things we all carry. Yes. These common variants cannot be sledgehammers, because if they were, they would be heavily weeded out by natural selection very quickly. That makes sense. Evolution wouldn't tolerate a sledgehammer in 30% of the population.
32:12Right. Instead, they are thtle nudges. And what the data shows is that for these common, subtle nudges that drive population level traits like height, intelligence, or psychiatric risk, the action is primarily in the regulatory formatting.
32:27Subtle nudges versus sledgehammers. I like that. Furthermore, even with the shift toward energetic regions, the authors note that for most phenotypes, the kudogenic regions, which is the combination of the exxons and the introns considered together as one unit, still contribute more overall heritability than the energetic regions.
32:46Wait, really? The genes plus introns still outweigh the energenic desards. Yes, for most phenotypes. The genes are still the anchor of biology. This paper is about where the heritability sits on a population level.
32:58It is not a claim that genes don't cause disease. That distinction between the sledgehammer and the subtle nudges is crucial. It really forces us to rethink how we categorize human biology from an evolutionary perspective.
33:09Basic physiological traits, how your body regulates lipids, how it processes sugar, how it builds. These are highly constrained, right? Extremely constrained. The basic mechanics of keeping a mammalian body alive, haven't changed much in millions of years, so they rely much more heavily on proximal coding variants.
33:27You can't mess with the formatting too much or the animal just dies. Exactly right. Basic somatic functions are heavily protected by negative selection. But when you move to higher order traits, like human cognition, personality traits, or psychiatric vulnerability, the architectural rule book totally changes.
33:44It does. These traits are relatively new in evolutionary history, and they are mediated by highly dispersed, distal regulatory elements that act across many different cell types. particularly in the central nervous system.
33:57Which brings us to a massive practical application of this research. Because this isn't just theoretical biology. This dictates how we must study and diagnose diseases moving forward in a clinical setting.
34:08It has massive clinical implications. Currently, if a patient goes in for clinical genetic testing. The standard of care, often for cost reasons, relies heavily on something called whole XM sequencing, or WES.
34:20Yeah, WES is incredibly cost effective because it intentionally ignores the dark matter. It only sequences that 2.55% of the genome that actually codes for proteins. Right, the lamp post. Exactly. If you were a clinical geneticist looking for a rare single gene sledgehammer mutation that causes a pediatric metabolic disorder, WES is a fantastic and efficient tool.
34:43But if we apply the findings of this paper to that clinical reality, we hit a wall. I mean, if you are a researcher trying to understand the genetic architecture of a highly polygenic behavioral trait or a psychiatrist trying to eventually develop genetic risk scores for schizophrenia.
34:59A standard array or hole like some sequencing will miss the bulk of this story. completely. It is functionally blind to the reality of the biology. If you are studying a psychiatric disorder and the Exxons only carry 13% of the heritability, WES is leaving 87% of the genetic risk completely unexamined.
35:16You are trying to read a book by only looking at the punctuation marks. That's a great way to think about it. To capture that energetic dark matter, to map those distal regulatory switches and understand the true architecture of complex disease, you absolutely must use whole genome sequencing or WGS.
35:33So, what does this all mean for pharmacology? If psychiatric risk isn't a single broken gear, but rather scattered across thousands and thousands of non-coding regions, constantly shifting depending on cell type, development stage, and environmental triggers.
35:48Does this make finding targeted treatments nearly impossible? A fair question. I mean, we can't design a pharmaceutical drug that targets 10,000 different energetic switches all at once. It is a daunting prospect, and it certainly explains why psychiatric drug development has historically been so incredibly difficult and prone to failure.
36:06We have been trying to use a single wrench to thook a malfunctioning neural network. So what's the alternative? Well, it doesn't make treatment impossible. It just means we have to stop thinking linearly and start thinking dynamically.
36:18We had to shift to a paradigm of network medicine. Network medicine. Explain how that works when the targets are so diffused. If psychiatric risk is mediated by thousands of tiny regulatory tweaks scattered across the non-coding genome, the biological reality is that those tweaks are not acting randomly.
36:36They're ultimately converging on shared downstream biological pathways. Oh, so they all funnel into a few main roads eventually. Exactly. For instance, thousands of different variants might all subtly affect the pathway responsible for synaptic plasticity, how neurons connect and communicate.
36:53Or they might converge on the brain's immune response, or on specific neurodevelopmental windows during adolescence. Okay, I follow you. The goal of future pharmacology won't be to fix a single broken, energetic switch.
37:05The goal will be to identify the hub where all these signals converge and develop a compound that nudges the entire network back into a state of equilibrium. It's like realizing you can't fix a traffic jam by talking to every single driver on the highway.
37:18You have to step back, look at the entire grid, and maybe just change the timing of one major traffic light to restore the flow. That is the promise of network medicine right there. And this paper provides the foundational map of where those regulatory signals are actually originating.
37:33It is an incredibly beautiful way to view biology. But as always, on this deep dive, we have to remain grounded in the scientific method, we have to look at the limitations of where the science actually is right now, today.
37:47Absolutely essential. This study is incredibly powerful. It's a monumental achievement in statistical genetics, but it's not the final word. Where are the blind spots in this paper? The authors are very transparent about their limitations, which is always the mark of rigorous science.
38:02First, while broad, the study only covers 34 traits and 74 functional annotations. Which is a lot, but not everything. Right. Right. Human biology is vast, and there are many more nuanced phenotypes and subcellular annotations that are not captured in this model at all.
38:18Makes sense. But the most crucial limitation. The one that the entire field of genomics must reckon with is the data source itself. This study relies heavily on European ancestry reference panels. Ah, yes.
38:31Let's expand on this because it's a systemic issue. They used the 1000 genomes phase 3 reference panel, specifically filtering for individuals of European ancestry. Why is it scientifically problematic to extrapolate these findings to the entire globe?
38:48It goes back to the concept of linkage to equilibrium that we discussed earlier, the neighborhood effect. Right, the chunks of DNA passed down together. The way DNA is chunked and inherited in neighborhoods is not universal.
39:00Because different human populations have different evolutionary histories, different migration patterns and different demographic bottlenecks. The structure of those genetic neighborhoods, the patterns of linkage to equilibrium and the frequencies of specific illes, can vary significantly between populations.
39:16So a statistical flag that points to a specific regulatory zone in a European population might not point to the exact same zone in an East Asian or an African population. Exactly, simply because the background architecture the DNA blocks is different.
39:29Oh, wow. So the map itself shifts depending on ancestry. The functional biology, how a cell actually works, is universal. But the statistical map to find those functions can be population specific. I see.
39:42Therefore, while the overarching principle that polygenicity drives heritability into the non-coding regions is likely a fundamental rule of human biology. The specific numbers and the specific genomic coordinates must be replicated across diverse ancestries to be considered truly universal.
39:58Because if we don't do that, If we only build our understanding of the non-coding genome based on one population, we risk severely exacerbating health disparities as we move into the era of precision medicine, because our diagnostic tools simply won't work as well for non-European populations.
40:15It's a crucial reminder that the map is not the territory, and right now, our map is missing a massive portion of the global population. We have a lot of work left to do. We do. But even with that limitation, the fundamental insight derived from this mixer-based framework is undeniably game changing.
40:31We have proven that the genetic architecture of our traits exists on a sliding scale. A beautiful dynamic sliding scale. Yeah. The more complex and polygenic a trait is, the more its heritability shifts away from the simple protein coating words of the exxons, and out into the dispersed, non-coding regulatory formatting of our genome.
40:52It is a humbling realization, honestly. The things that make us most complex, the treats that define human cognition and vulnerability are governed by the spaces between the genes, acting together as a massive, exquisitely orchestrated network.
41:05So I leave you, the listener, with this thought to mol over. What does this mean for the future of personalized medicine when the instructions for our most complex human traits aren't found in the tangible genes themselves, but remain hidden, dynamically interacting in the spaces between them?
41:19It's going to be a fascinating future. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow, or subscribe in your podcast app and leave a 5 star rating.
41:36If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
41:46Thanks for listening and join us next time as we explore more science base by base.