This paper introduces Epi-PRS, a workflow that uses genomic large language models to impute cell-type-specific epigenomic features from diploid genotypes and trains nonlinear risk models to improve polygenic prediction from WGS. The method improves AUC for breast cancer and type 2 diabetes in UK Biobank and shows gains from modeling regulatory context and rare variants.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Glad be here. So um, I have a question for you.
0:11What if having your entire genome sequenced isn't actually enough to accurately predict your risk for complex diseases? Oh, that is the big question right now. Right. Because for years, we've treated DNA like, well, like this simple, straightforward lot of ingredients.
0:29Yeah exactly. I mean, if you've ever spit in a tube for one of those consumer genetics test, you've essentially bought that ingredient list. Which is cool, but it's limited. Right. We sequence your genome.
0:40We read those raw components and we just assume we know exactly what kind of cake is going to come out of the oven. But baking is way more complicated than just having flour and sugar. Exactly. So what if the real secret lies in understanding how those ingredients interact?
0:54What if the key to your health is really about how those instructions are actually, you know, cooked by the body? That's where things are heading. Because 1000000s of genomes have been sequenced around the world.
1:04Yet our predictions for conditions like breast cancer or diabetes are still hitting this stubborn ceiling. It's incredibly frustrating for researchers. I bet. It really makes you wonder, how could this change our understanding of our own genetic blueprints?
1:18What really happens when we start reading DNA, not just for its raw code, but for its regulatory context? Well, today we celebrate the work of the team at Stanford University. Specifically, Wan Wen Zheng, Hanman Guao, Kyolu, and Wing Hong Wong, who have advanced our understanding of polygenic risk prediction.
1:36It some truly groundbreaking work. It really is. And their breakthrough findings or detail in a paper titled improving polygenic prediction from whole genome sequencing data by leveraging predicted epigenomic features.
1:48Yeah, very scientific. It was published in PNAS on June 12, 2025. Okay, let's unpack this because for anyone who has been, you know, following genetics. The Holy Grail for a while now has been the polygenic risk score.
2:02or PRS. Right, the PRS. We hear about these scores all the time. But why is the traditional way we've been calculating them, suddenly hitting this invisible wall? That's a really good question. Especially now that we have these massive whole genome sequencing databases like the UK biobank, which holds data on half a 1000000 participants.
2:24It's an unbelievable amount of data. Right. So with all that deep rich data. Shouldn't our predictive models be practically perfect by now? Well, you certainly think so, but it really comes down to the architecture.
2:36Specifically how traditional PRS methods like LDPred 2 and PRSCS were fundamentally designed. Oh, so the old models just weren't built for this kind of data. Exactly. They were built for an older era, primarily from S&P arrays, which mostly just capture common genetic variants.
2:51But more importantly, these traditional methods rely on strict linear mathematical models. Linear, meaning what, exactly? Well, a linear model assumes that the effect of every genetic variant simply adds up in a vacuum.
3:03So variant A adds a 2% risk, variant B adds a 3% risk, and you just tally them together on a ledger. Right. So going back to that baking metaphor earlier, the linear model is essentially just dumping flour, sugar, and raw eggs onto the kitchen counter.
3:19Yes. And then just looking at the pile and calling it a cake. That is a perfect way to conceptualize it. It completely ignores the actual process of how those things mix and interact. Because biology isn't just a simple math equation.
3:32Biology is inherently nonlinear. The ingredients don't just sit next to each other, they react. And now, hold genome sequencing has just blown the doors wide open, it doesn't just look at those common variants anymore.
3:44It sees everything. Yeah, it detects rare and even de novo, meaning completely new non-coding variants. Wow. And traditional linear models completely ignore how these rare variants might interact with each other in complex ways.
3:57So they're just blind to it. Completely. Furthermore, and this is really the crucial missing piece that this whole deep dive is about. They entirely miss the epigenomic layer. The epigenomic layer. Okay, so if the DNA sequence is the list of ingredients, the epigenomics would be like the oven temperature.
4:14Yes, the temperature, the mixing speed, the actual timing of the bake. Wow, okay. Precisely. We are talking about the regulatory mechanisms, elements like enhancers, promoters, and uh, transcription factor, binding sites.
4:29These are the things that actually control the genes. Exactly. They are the hidden switches that actually control gene expression. They decide if a gene is turned on, turned off, dialed up to a 10 or dialed down to a two.
4:40That makes a huge difference. It really does. If you have a genetic variant that alters one of these regulatory switches, it can have a massive cascading impact on your disease risk. But the old bottles don't see that.
4:52Right. Traditional PRS methods either completely ignore this regulatory traffic or they just try to forcefully cram it into that same old additive linear box. Which brings us to the innovation from the Stanford team.
5:04Yes. They developed a completely novel framework called EpiPRS to tackle this exact blind spot. And it's a brilliant approach. And thinking about this. If traditional PRS is like looking at a static 2D map of a city's roads, EPRS is like having a real-time AI traffic predictor.
5:23Oh, I like that analogy. Right. Because it knows exactly how a roadblock on a tiny side street affects the flow of the entire highway system. It sees the whole interconnected mess. It absolutely does. But I have to push back here.
5:36How does an AI actually do this with DNA? I mean, we're taking a static sequence of A, Cs, Ts, and Gs? How does it look at that and predict the real-time dynamic traffic of the human genome? That is the core of their methodology, and it's honestly a brilliant piece of computational engineering.
5:52Okay, walk me through it. So the team broke EpiPRS down into 3 major steps. First, they perform what's called personal genome construction. Personal genome construction. Yeah. Using computational tools, like beagle and VCF 2 deployed, They don't just look at a generic flattened sequence.
6:09What do they do instead? They actually reconstruct both the maternal and paternal genomes for an individual. They build your specific 2 lane deployed highway. Oh, because you get one set of instructions from your mom and one from your dad.
6:22Exactly. And just like in real life, those instructions might be entirely contradicting each other. Like, mom's DNA might be yelling to turn out the heat while dad's DNA is trying to pull the cake out of the oven entirely.
6:33That is exactly it. And you absolutely need both to see those true interactions. You have to know who is winning that argument at the cellular level. makes total sense. So what's step two? Once they have that personalized deployed genome, they move to step two, which is epigenomic feature extraction.
6:50Okay. This is where the AI truly enters the picture. They use a genomic large language model or LLM called informer. Wait, okay, stop right there. When most of us hear large language model, We think of chat GPT, writing emails or generating code.
7:06Yeah. How does an LLM work on biology? I mean, DNA is in English. It is in English, sure. But DNA is a language. How so? Well, it has a 4 letter alphabet, and those letters form words and sentences that the cell actually reads.
7:20Oh, just like chat GPT ingested massive amounts of human text to learn the rules of grammar and, you know, predict the next word in a sentence, and former ingested massive amounts of genomic data. So it's predicting the next word of DNA.
7:34Not quite. Instead of predicting words, it learned the grammar of how DNA structural motifs, fold, bend, and interact in three-dimensional space. That is mind blowing. It really is. For a given input sequence of 196,608 base pairs.
7:50Informer predicts 5313 different epigenomic signals. Wait, it's predicting 1000s of different ways the DNA might be accessed or modified just from reading the raw text. Yes, it is predicting chromatin accessibility, histone modifications, transcription factor binding the entire regulatory landscape.
8:08Incredible. It literally looks at the text and visualizes the physical 3D architecture of the genome. But practically speaking, predicting 5,313 signals over nearly 200,000 base pairs. That is a staggering wall of data.
8:23It's huge computational challenge. If you feed all of that raw information into a predictive model, it doesn't just choke. Like, how does it not get completely overwhelmed by the noise? You've hit on a major hurdle there, which actually leads us directly to their 3rd step.
8:36Because you're right, that much data would cause a standard model to overfit. Overfit, meaning. Meaning it would just memorize the random noise instead of learning the actual patterns. Oh gotcha. So they use a mathematical technique called local PCA or principal component analysis to reduce the dimensions of the data.
8:55So for those of us not deep into data science, PCA is basically like taking a dense 500 page technical manual and perfectly summarizing the most crucial points. Yeah. Like distilling it down to a 10 page exec. Correct.
9:08Think of GBRT as a committee of decision trees that learn from their mistakes. A committee? Yeah. The 1st tree makes a prediction. realizes where it was wrong, and the next tree specifically focuses on fixing those errors.
9:21That's really clever. Because it's a tree-based model, it naturally captures those complex, nonlinear interactions we talked about earlier. So it connects all those interacting factors. Exactly. What's fascinating here is that EPPRS uses these imputed epigenomic signals as intermediaries between the raw genotype and the final physical trait.
9:42It bridges the gap. It does. It's not just hunting for a statistical correlation in a vacuum. It's mapping a biologically meaningful pathway. It's explicitly saying, this rare genetic variant changes this specific regulatory switch, which changes how this gene is expressed, which then alters your disease risk.
10:01So it's finally connecting the dots between the raw code and the physical reality of the human body. Yes. That sounds incredible in theory, but simulations are great in this sterile lab setting. Human biology is notoriously messy.
10:13Very messy. When they threw actual data at this. you know, real people with real complex environments. Did the AI actually hold up? To test that, they started with incredibly rigorous simulations based on the real genetic architecture of the UK biopank.
10:29Okay, so using real world blueprints. Right. They wanted to see exactly what happened when you introduce rare variants into the mix. What are they finding? They found that in data sets with over 2000 instances, the nonlinear models heavily outperformed the traditional linear models. But the real revelation came when they simulated a scenario where up to 75% of the variants affecting a trait were rare variants.
10:52Oh wow, 75%. Yes. In that environment, FEPRS increased the predictive accuracy by up to 20% compared to traditional methods like LDPred 2 and PRSCS. Okay, let's pause there. Because a 20% jump is massive when we're talking about predicting disease risk.
11:08It's huge. And here's where it gets really interesting. Taking this AI traffic predictor out of the simulated lab and applying it directly to real-world UK biobank patients. Right, the real test. How did it handle actual messy human records?
11:22It was incredibly successful. They focused on 2 major, highly complex diseases, breast cancer and type 2 diabetes. To incredibly impactful conditions. Absolutely. For breast cancer, they tested the model on 10,547 actual cases and an equal number of controls.
11:40That's huge sample size. It is. They look specifically across 5 linkage to equilibrium or LD blocks. You can think of LD blocks basically as neighborhoods of DNA where we know disease signals strongly cluster together.
11:53Okay, DNA neighborhoods. like that. And EPRS vastly outperformed the traditional whole genome methods. How did they measure the performance? To measure this, they used a relative performance score called Lambda.
12:050 is a complete random guess, and one is a flawless, perfect prediction. Okay, got the scale. So how did the old models do? The traditional methods maxed out at a Lambda around 0.27, EPPRS achieved almost 0.40.
12:18I want to push back on that for a 2nd because context is everything. A Lambda .40 sounds impressive mathematically compared to .27. But in the grand scheme of predicting a disease is notoriously complex and devastating as breast cancer.
12:33What does that actually mean for a patient? Is this moving the needle clinically, or is it just a nicer looking graph in a research paper? Oh, it absolutely moves the clinical needle. Think about how polygenic risk scores are used in practice.
12:45If someone is in the 40th percentile of risk under the old linear model, they might be told they are at average risk. Right. They're told not to worry, just follow standard screening guidelines. Exactly.
12:57But what if EPRS by capturing all these hidden nonlinear regulatory interactions reveals that they are actually in the 95th percentile of risk? Oh, wow, that changes everything. Their clinical care changes entirely.
13:11They get screened earlier, more frequently and with more sensitive tools. So that mathematical jump is actually catching high risk patients who are slipping through the cracks. Exactly. That jump from .27 to .40 represents 1000s of hidden, high risk individuals suddenly becoming visible to the medical system.
13:28Wow, okay. That is a tangible life-saving difference. It truly is. What about the type 2 diabetes data? The results for type 2 diabetes were even more dramatic. They use 20,000 cases and 20,000 controls across 11 of those DNA neighborhoods, those LD blocks.
13:44Okay. The baseline models scored lambda values anywhere between .077 and .323. Not great. Not great at all, but EPPRS achieved a lambda of .515. Wow. It absolutely crushed the baseline models. So it's not just a minor incremental upgrade.
14:01It's a fundamental leap in how accurately we can foresee these conditions. It is, but perhaps the most satisfying part of this entire study is a deeply biological validation that emerged organically from the data.
14:13What do you mean organically? Well, when the researchers opened up the AI's black box to analyze the top 50 predictive features that it used to determine type 2 diabetes risk. To see how it was actually doing the math.
14:24Right. They found that 44% of those signals originated specifically from pancreas or liver tissues. Wait, to weave our analogy back in here. So the AI didn't just find a traffic jam. It looked at the map and realized that the main bridges and the pancreas and liver were the most critical choke points for diabetes.
14:42And it did that without the researchers ever telling it to look there. Exactly. They didn't explicitly program the model to say, hey, focus on the pancreas. diabetes happened. That's incredible. The AI successfully identified the most disease relevant cellular context entirely on its own.
14:59Wow. This proves that the model is biologically meaningful, not just a mathematical parlor trick. You actually learn something real. Yeah, it placed significant weight on the active enhancers and promoters in the exact tissues we know clinically are responsible for the disease.
15:14Just by reading the structural grammar of the DNA. Exactly. So what does this all mean? We have this incredibly powerful tool that can take a person's whole genome sequence, run it through a genomic language model, figure out the hidden regulatory traffic patterns, and predict disease risk with unprecedented accuracy.
15:34That's a massive step forward. But does this mean the problem is completely solved? Because I have a critical question here about how this thing is actually built? How for it. If this EpiPRS model? is so heavily focused on the regulatory side.
15:48The epigenomic features, the enhancers, the promoters. Doesn't it kind of have a massive blind spot? So? Well, what happens if you have a mutation that doesn't mess with the regulatory switch? but directly breaks the proteins code itself.
16:03Doesn't the model miss that entirely? This raises an important question, and you've hit on the exact limitation that the authors themselves point out very clearly. Well, they caught that too. They did, because EPRS relies entirely on informer annotations.
16:16It is beautifully designed to capture transcriptional right. the switches and dials, but it might completely miss genetic effects that act through direct changes to the protein coding sequence itself or through RNA processing or translation.
16:30So if a variant break the actual gear rather than the switch, EPPRS might not see it. So it's like our AI traffic predictor can flawlessly tell you when a traffic light is broken and causing a backup. But it won't notice if the bridge itself has physically collapsed.
16:46That's a great way to ground the analogy. The researchers note that moving forward, the ultimate predictive model will likely need to be a hybrid. Combining both approaches. Right. They will need to combine direct genetic features to catch those broken bridges.
17:00With these predicted epigenomic features to catch the broken traffic lights. That makes a lot of sense. They also highlight some really exciting future directions for the technology itself. For instance, informer is incredibly powerful, but it's computationally heavy.
17:14can imagine. So future versions could use faster, highly optimized, large language models, like one called Kukitay to dramatically speed up inference. Which would be essential if you want to eventually roll this out to 1000000s of patients in a standard hospital setting, right?
17:30Rather than just running it on a supercomputer at Stanford. Exactly. It needs to be scalable. They also discussed using higher resolution models like BPNet, which can predict transcription factor binding down to a single base pair resolution.
17:44Single base pair. That's an astonishing level of detail. It is. And crucially, they plan to explicitly incorporate real world confounders into the model, like geographical population structures. Lies are that important?
17:57Because right now, the vast majority of genomic data comes from populations of European ancestry. Oh, right. That bias in the data. Yeah, to make sure these predictions remain unbiased and equally effective across all diverse demographics, those geographic and ancestral structures need be woven into the model's architecture.
18:16Well, it sounds like we are truly just scratching the surface of what genomic LLMs are going to be capable of. We really are. It's wild to think that the same underlying architectural concepts that write our emails right now are being used to decode the physical regulatory landscape of human disease.
18:30It truly is a new frontier. If we were to synthesize all of this deep dive into a core takeaway. Please do. EpiPRS represents a paradigm shift by utilizing large language models to translate raw personal DNA sequences into predictive tissue specific epigenomic features.
18:48By successfully capturing the complex, nonlinear interactions of both common and rare genetic variants, it dramatically improves the accuracy of polygenic risk scores over traditional linear models. What does this mean for the future of personalized medicine, and how we might one day map not just our genetic fate, but the hidden regulatory switches that control it?
19:09It's an exciting time for science. That's for sure. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
19:21If you enjoyed this, follow or subscribe in your podcast app, and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
19:34Thanks for listening, and join us next time as we explore more science, base by base.