TWAS using genetically predicted expression exhibit polygenicity-driven inflation that increases with GWAS sample size and heritability; a gene-specific variance-control correction yields calibrated p values.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So, um, I want you to imagine for a 2nd that you are tasked with finding a specific criminal in a massive crowded city.
0:17Okay, setting the scene. Right. And you are armed with this state of the art facial recognition algorithm. When you test this tool in a controlled environment, say pointing it at a group of 10 or maybe 20 people, it works flawlessly.
0:30Yeah, it spots the target without any hesitation. Exactly. But then you deploy that exact same algorithm on high dimensional data. You pointed at a stadium crowd of like 100,000 people. And suddenly the system doesn't just struggle.
0:43It completely panics. It gets overwhelmed by the sheer volume of overlapping features and background variables, and it just starts throwing red flags everywhere. It ends up identifying practically everyone in the stadium as a suspect.
0:55Right, which is a profound failure scale. And, well, it perfectly captures a systemic crisis that's currently unfolding in modern computational genetics. Yeah, because we are seeing this exact kind of hallucination in our genetic databases right now.
1:10When researchers use our most advanced statistical tools to find the specific causal genes driving complex diseases, the software is essentially screaming false positive. Just constantly overreporting.
1:22Yes. And it's doing it simply because the crowd of genetic data, the statistical background noise has just gotten too large for the baseline models to handle. So what really happens when the very analytical frameworks we rely on to understand complex diseases start hallucinating connections that simply aren't there.
1:40It's massive problem. It really is. And, you know, how could correcting this mathematical artifact fundamentally change the future of drug discovery? That is what we're getting into in this deep dive. Absolutely.
1:51And today we celebrate the work of Yan Yulian, Festus Nisimi, Haikung IM, and their entire team at the University of Chicago, who have advanced our understanding of genetic association studies. Yes, a huge shout out to them They've tackled this exact problem head on.
2:06Really fundamentally changing how we interpret massive genetic data sets. They identified, and more importantly, corrected, a major source of statistical inflation. And I think it is totally worth noting, the sheer scale of the computational resources that made this breakthrough possible.
2:24I mean, we were talking about critical support from the Argon National Laboratory. Oh, definitely. You need that kind of computing power. Right. And the utilization of the massive UK biobank, having this repository of genetic and health data from 100s of 1000s of individuals, well, it provides that ultimate stadium crowd we talked about to test these algorithms against.
2:44The scale is absolutely what makes the analysis possible. But, you know, to understand why our standard tools are failing in that crowd, we 1st need to critically examine what transcript on wide association studies or to us are actually trying to achieve in the 1st place.
2:57Okay, let's unpack that. Because we rely heavily on Twas to map G-Was low site to gene expression. We are essentially looking for the middleman, right? Exactly. The biological mediators. Right. So we know a genetic region is associated with a disease, but we want to know if the expression level of a specific gene in that region is the actual lever being pulled.
3:18The problem is, researchers have been noticing massive inflation in their Tijua's results, like way too many genes are being flagged as statistically significant. Yeah, and for a really long time. The field kind of misdiagnosed the cause of that inflation.
3:34Wait, really? blaming it on? Well, the prevailing assumption was that our predictive models for gene expression, which are often drained on reference panels like G Tatex, were simply too noisy. Ah, the logic was, you know, that inaccurate expression weights were feeding bad data into the association tests, and that was generating the false positives.
3:54Okay, I actually need to push back on that a little bit. Because if we look at fundamental statistical principles, specifically error and variables theory, why wouldn't researchers just assume their Tiwa's models were underpowered?
4:05That is the exact right question to ask. Right. Because having a noisy predictor variable usually increases your standard error, it drops your Z score, it doesn't spontaneously fabricate statistical significance out of thin air.
4:19Yes. And that is the fundamental contradiction that the U Chicago team recognized. Blaming the proxy model is intuitive, but statistically, it is entirely backwards. A noisy predictive model blurs the signal.
4:33It makes it harder to cross the significant threshold, not easier. So what was actually generating all these false signals? The team realized the real culprit wasn't random noise in the expression models.
4:44It was the widespread, correlated genetic architecture of the diseases themselves. Meaning the polyogenic background of the trait? Exactly. I mean, most complex traits, like height, heart disease or psychiatric disorders.
4:56They aren't controlled by one single gene. They're driven by 1000s of small effect variants. It's everywhere. Yeah, so the genetic predictors we use for gene expression inadvertently capture some of the variants driving the disease itself.
5:07It's not random noise. It structured biological confounding. Oh, wow. So the model detects a correlation and just assumes causation. Exactly. It flags the gene as a mediator. When it is really just a bystander caught in the polygenic crossfire.
5:22Okay. I mean, that makes total biological sense. But proving it computationally seems incredibly difficult. If real human traits are inherently messy and deeply polygenic, how did they isolate this, this mathematical artifact?
5:38A great question. Because you can't just take a real complex trait and magically erase the causal genes to see if the software still hallucinates, right? No, you can't do it biologically. But you can do it computationally.
5:49And their methodology for proving this was just exceptionally elegant. They decided to simulate the void, basically, by constructing a polygenic null trade. Okay, walk me through how they built that synthetic phenotype because that sounds fascinating.
6:03So they utilized real human genotype data from up to 100,000 individuals in the UK biobank. Which is key, right. They didn't just generate random numbers. Exactly. This is crucial because using real DNA preserves the complex real world linkage disequilibrium structure of the human genome, it keeps all the messy correlations intact.
6:25Right. But instead of analyzing a real clinical phenotype like diabetes or hypertension, they mathematically generated a fake disease. A completely fake disease. Yes, and they program this synthetic trait to have a highly complex polygenic architecture, meaning it was influenced by 1000s of random genetic variants across the genome.
6:45Okay. However, they explicitly engineered it so that the true causal effect of the target genes expression on this trait was absolute zero. Ah So the trap is set. The trap is perfectly set. They have a purely polygenic phenotype with 0 actual connection to the gene expression being tested.
7:01Which means if standard TWS software, like predicts can or fusion, flags any gene as significantly associated with this fake disease, we know with mathematical certainty that it is a false positive. Exactly.
7:14And the results were stark. The standard tools failed spectacularly. They threw massive numbers of false positives. I bet. And what makes this study so robust is that they didn't limit this stress test to just gene expression.
7:25They wanted to determine if this was an algorithm of vulnerability inherent to any high-dimensional association study. So they broadened the scope. They did. They applied this same null trait framework to predict 580 different metabolite levels and 471 brain MRI phenotypes.
7:45Wow. Just to see if the statistical hallucination persisted across entirely different types of biology. And it did. The inflation was universal across the transcriptums, the metabolome, and the neuroimaging data.
7:57But more importantly, the failure wasn't random. What do you mean? The team identified a strict, predictable mathematical pattern to the hallucinations. The rate of false positives scaled linearly based on 2 specific parameters.
8:10Okay, let's look at the math there, because this is where the paper really shifts from diagnosing a problem to actually solving it. What were the 2 parameters driving this linear explosion? The first is the sample size of the GWS cohort, which we denote is in.
8:23And the 2nd is the local heritability of the trade, specifically denoted as h squared delta. Let me just clarify that 2nd parameter real quick for the listeners. We aren't talking about the global heritability of the entire disease, right?
8:36We're talking about the local heritability, um, the proportion of phenotypic variants explained by the specific genetic region being tested. That is a critical distinction, yes. It is the local genetic architecture driving the regional confounding.
8:50Yeah. But let's look at the implications of that 1st parameter. sample size. Oh, this highlights a massive paradox for the genomics community. I mean, the entire trajectory of modern genetics is pushing toward larger data sets.
9:02Right, bigger is better. Exactly. We want bigger biobanks, multi-ancestry cohorts, 1000000s of participants, because the assumption is that higher end equals greater statistical power and clarity. Right. But this linear equation dictates that as our sample sizes grow, the false positive rate scales linearly right alongside them.
9:21More data is literally generating more hallucinations. It is the terrifying irony of big data in this specific context. If you are studying a highly polygenic trait and you double your GOEA sample size, you are linearly amplifying the background, confounding that the algorithm misinterprets as a true causal signal.
9:41That's wild. It is. But because this error is perfectly linear and predictable, it means it can be mathematically corrected. Right. If you can measure the inflation, you can reverse engineer it. Exactly.
9:53So the researchers quantified this baseline inflation rate for every single gene, creating a parameter they called the fi factor, represented by the Greek letter fi. And they found that for most genes, this inflation slope sits roughly around 10 to the -5th.
10:08Okay, 10 to the negative fit. That sounds like a vanishingly small number. On its own, it is, but when you multiply 10 to the -5th by a modern GWE sample size of half a 1000000 individuals, it creates a massive distortion in your Z scores.
10:20Right, because the end is just so huge. Precisely. So to implement the fix, they developed a variance control method. You take the raw inflated Z score from your standard association software, and you divide it by a specific scaling factor.
10:34And that formula is the square root of the quantity one +5 xen times local heritability. Yep. So square root of one +5 times N times H squared delta. You got it. By dividing the raw statistic by that scaling factor, you systematically shrink the hallucinated results back to a standard normal distribution.
10:54Oh I see. You recalibrate the test, so the false positive rate drops back to the expected baseline without aggressively destroying the true biological signals. That is so elegant. We really need to talk about what happened when they applied this variance control method retroactively because they took this fix and ran it across 110 real world GWS traits.
11:14Yeah, the real world applications where it shines. The before and after data is staggering, particularly for highly polygenic traits like psychiatric disorders. I mean, schizophrenia, bipolar disorder.
11:24These are notoriously complex architectures. The correction was profound. Before the mathematical fix, the standard TW's software was flagging a median of 12 significant genetic low sci for these highly polyogenic traits.
11:3712. Okay. After applying the variance control method, that median plummeted down to four. Wow. A two thirds reduction in, quote unquote, significant findings, which forces us to look at the broader implications here.
11:50Specifically regarding pharmaceutical economics and drug discovery, because target validation is arguably the single biggest bottleneck in developing new therapeutics. Oh, without a doubt, developing a novel drug costs 1000000000s of dollars and takes over a decade of clinical trial.
12:05And pharmaceutical companies increasingly rely on human genetic evidence to select which biological targets to pursue. If a company bases a massive phase 2 clinical trial on one of those 8 false positive genes, if they target a gene that is just a bystander to the polyogenic background, that drug is guaranteed to fail, it will not modulate the disease mechanism.
12:25So by implementing this variance control, we aren't just cleaning up spreadsheets. We are literally preventing drug developers from chasing biochemical ghosts. Exactly. We are saving years of wasted research and reallocating resources toward true causal mediators.
12:43So it's almost like noise canceling headphones. Oh, I like that. Walk me through how you're seeing that. Well, think about being on an airplane. The widespread polygenicity of the trait is the roar of the jet engine.
12:54It's loud, it's everywhere, and it permeates the whole environment. Our old standard models were essentially just turning up the volume on everything to try and hear the person sitting next to us, which just made the engine roar absolutely deafening.
13:07But this variance control method. Mathematically isolates the specific frequency of that background engine, which is the 5 factor and plays the exact inverse computational wave to cancel it out. Leaving just the true causal signal intact, the actual conversation right next to you.
13:22That's a great way to look at it. It isolates the true mediator from the ambient genetic noise. It is a brilliant conceptualization. But you know, as we evaluate any major methodological shift in genomics, We do have to critically examine the limitations of the model.
13:37Of course, nothing is perfect. Right. This variants control approach relies heavily on the infinitesimal model assumption. Meaning it assumes the genetic architecture of the trait consists of thousands of tiny, relatively evenly distributed effect sizes.
13:52Correct. And for most complex human traits, the infinitesimal model is a very robust approximation. However, if a researcher is studying a disease with a highly unusual genetic architecture, say an oligogenic trait driven by a handful of massive lowsi rather than 1000s of small ones, this specific mathematical adjustment might overcorrect or behave unpredictably.
14:14Furthermore, while this method is exceptional at neutralizing polygenic background noise, it does not resolve the issue of local pleotropy. Okay, let's distinguish between those two because they often get conflated.
14:26Polygenic noise is the collective genome wide background confounding. What is local playodropy in this specific context? Local pleotropy occurs when a single genetic variant independently drives 2 completely separate biological outcomes.
14:41So a specific locus might increase the expression of a certain gene, and through a completely divergent biological pathway, that same locus increases your risk for a disease. Ah, so the TWS software looks at that coordinate and flags the gene expression as the cause of the disease, when in reality, they are just parallel effects stemming from the exact same genomic address.
15:04Exactly. And the variance control method will still flag that relationship as significant, because the statistical association is genuinely real. In the case of local pleiotropy, we aren't dealing with a mathematical hallucination driven by sample size, we're dealing with a biological confounder.
15:19Right. So to untangle that, researchers still need to rely on sophisticated colloquialization tools to determine if the causal variants for the expression and the disease are truly shared, or if they're just physically adjacent to each other.
15:31Precisely. Variance control is not a panacea for all genetic confoundings. But by clearing away the massive systemic cloud of polygenic false positives first, it dramatically reduces the search space. Which makes it much easier to deploy those computationally heavy colloquialization tools on the low side that actually matters.
15:50Exactly. It clears the field so we can see the true biological architecture. So to synthesize the core mechanics of the steep dive, as genetic association studies scale up in sample size, the widespread polygenicity of complex straits systematically inflates false positive results across transcriptomic, metabolomic, and imaging data.
16:10By implementing a novel, gene specific variance tintral method based on the 5 factor, researchers can mathematically neutralize this linear background noise, ensuring that the biological mediators we target are truly driving the disease.
16:23It really is a necessary recalibration of how we handle high-dimensional biological data. It is, but it leaves us with a highly provocative thought to consider as we look back at the literature. If we are only just now implementing this mathematical correction to filter out the background noise of big data.
16:42What does this mean for the 1000s of previously published genetic associations that didn't use this variance control? This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
16:58If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
17:10Thanks for listening, and join us next time as we explore more science base by base.