A concise review of how post-GWAS methods are being used to move from statistical associations to translational insights by integrating drug-target prioritization, single-cell resolution of regulatory mechanisms, and imaging-derived organ phenotypes.
0:00Welcome to Base by Base, the papercast that brings genomics to you, wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine, just for a 2nd, handing a pharmaceutical company $2000000000 and like 15 years of your life, only to flip a coin at the finish line to see if the medicine they created actually works.
0:20Yeah, it's I mean, it sounds absurd, but that is the terrifying reality of modern drug development. Right. It' just a staggering problem. And honestly, that coin flip analogy is barely an exaggeration.
0:32Barely. Because even after over a decade of research, all those 1000000000s invested. You know, nearly 50% of drugs fail in the late stages of clinical trials. And it's simply because they lack efficacy, they get into the human body and just they do not perform the biological task they were designed for.
0:50Which is wild. And then what, another 25% fail because of unforeseen safety concerns? Exactly, safety or toxicities that we just didn't see coming in the animal models. So we have a system where basically 3 quarters of these massive medical investments evaporate right before reaching the patients who actually need them.
1:07So the question is, how could this massive failure rate change if we could, you know, read the exact instruction manual of human disease before we ever even tried to design a drug? Well, that potential to frontload our understanding, that is the absolute promise of modern genomics.
1:25We are moving past just identifying that a disease risk exists somewhere in your DNA. Right, like just knowing you have a risk. Yeah, and we're moving to actually decoding the physical, the cellular mechanics of how it operates.
1:38And today we celebrate the work of the team at the University of Cambridge, along with researchers from Harvard and the Baker Harden Diabetes Institute who have advanced our understanding of multi-scale genomics.
1:48It's really hard to believe it has been nearly 25 years since the initial draft of the human genome. I know, time flies. It really does. And for the past 2 decades, you know, the heavy lifter in this field has been genome wide association studies or gelo class.
2:02And from my reading, I mean, DuOS was incredible at feeding genetic low-si, right? Like these specific zip codes in our DNA that are linked to diseases. It was a massive breakthrough. Yeah. But we seem to hit this massive scientific bottleneck.
2:16Like we found the statistical associations, but we were completely missing the biological mechanism. And that bottleneck has been, um, well, really the defining frustration of genomics for years. GWOS gave us this massive list of genetic coordinates.
2:31It allowed us to say, you know, that people with a variation at coordinate X have a higher risk of heart disease. But just having the coordinate doesn't really help you fix it. Exactly. A coordinate on a map doesn't explain the biology.
2:42It doesn't tell you the how or the why. And without understanding the actual mechanical breakdown happening inside the body, trying to design a drug to fix it is, well, it's essentially guesswork. If I'm understanding this correctly.
2:55It's kind of like, um, GDOB was is basically a satellite map showing a traffic jam in a city. Okay, I like that. Right. Like it tells you exactly where the congestion is, but from space, you have no idea why the cars are actually stopped to figure out if it's, I don't know, a pothole, a broken traffic light, or a parade, you have to zoom all the way down to the street level.
3:16The street level view is a really great way to think about it. And achieving that view is the entire premise of this new multiskale genomics. We have to move away from just, you know, population level statistics and actually integrate data across 3 distinct biological scales.
3:32Which are the molecular level, the cellular level, and the whole organ level, right? Yes, exactly. So let's talk about how researchers actually pull off that zooming process. Because when we look at the cellular level, the sources talk heavily about single cell RNA sequencing.
3:47Yes, CRNA sec. It's revolutionary. Because our bodies aren't just one uniform block of DNA, right? We have skin cells, neurons, heart muscle cells. They all cheer the exact same genome. Right, but they operate completely differently.
4:01So this sequencing tech, it lets you see which specific genes are turned on or off in a single specific cell at any given moment. Exactly. It provides the immediate genetic context. But, you know, just observing that a gene is turned on doesn't actually prove that a specific genetic variant is controlling it.
4:18Oh, interesting. So how do you prove that? Well, to bridge that gap, researchers use something called expression quantitative trait LOSI mapping. Okay, EQTL map. Right, EQTL mapping. This technique links a genetic variant directly to the regulation of a gene within that specific cell type.
4:35So it proves a causal chain of command. Okay, so that handles the cells and the molecules. But scaling up to the whole organ level, that seems vastly more complicated. It is. It requires a massive leap in data integration.
4:47Which brings us to imaging genetics. Researchers are tapping into these population scale databases, most notably the UK biobank. Oh, I've heard of them. They just hit a big milestone, right? They did. They recently crossed a milestone of 100,000 MRI and PE spans of the brain, body, and bones from volunteers.
5:07Wow. 100,000. Yeah, it's a gold mine. And they feed this visual data into deep learning models to extract high-dimensional data points, which they call imaging drived phenotypes or IDPs. Wait, I wanna make sure I'm grasping this.
5:23They have an AI analyzing 10s of 1000s of medical scans, mapping out the microscopic physical shapes and structures of these organs, and then cross-referencing those physical shapes with the patient's genetic codes.
5:35That is the foundation of it, yes. But, I mean, generating the data is really only half the battle. The real innovation is the modeling used to interpret all of it. They rely on these frameworks like stratified linkage to equilibrium score regression.
5:47Okay, stop right there. Stratified linkage to equilibrium score regression. That sounds like someone dropped a dictionary into a blender. I know. Scientists are terrible at naming things. What is that actually doing in plain English?
5:58Fair enough. Think of it as an incredibly sophisticated calculator. It doesn't just find a broken gene. It calculates exactly how much that specific genetic variation contributes to a physical, structural change in a specific cell or organ.
6:16Oh, right. So it measures the heritability of a disease and maps it directly onto the physical architecture of the body. Which brings us back to those failing drug trials we started with. If we can map the exact architecture of a disease, it makes perfect sense that we can figure out which drugs are actually worth developing.
6:31Exactly. The pharmaceutical industry calls this, therapeutic target prioritization, and honestly, the numbers are definitive. Oh, definitive. Well, therapeutic targets supported by human genetic information are 2.6 times more likely to succeed in clinical trials.
6:452.6 times. Yeah. And that single metric represents 1000000000s of dollars in years of save time. I'd love to ground this in a concrete example. Like, how does knowing the genetic architecture translate into a physical drug on a pharmacy shelf?
6:58One of the most elegant examples of this comes from studying rare genetic mutations. Years ago, researchers noticed that a specific subset of people of African descent had unusually low cholesterol levels.
7:11And they had almost 0 risk of cardiovascular disease. It turned out they carried a rare natural loss of function mutation in a gene called PCS canine. So nature basically ran a biological experiment showing that if this specific gene is broken, people don't get high cholesterol.
7:29Exactly. But how does that actually work mechanically? Like, what is PCS canine doing that makes breaking it a good thing? The mechanism is really fascinating. In a normal system, the liver have these LDL receptors on its surface.
7:42Think of them like tiny vacuums pulling bad cholesterol out of the blood. Okay, vacuums, got it. The PCS canine proteins job is to bind to those vacuums and destroy them, so the liver doesn't clear too much cholesterol.
7:53But if you have this rare mutation, you don't produce the PCSK9 destroyer protein. Which means the liver gets to keep all of its cholesterol vacuums. They just stay on the surface of the cell, constantly pulling bad cholesterol out of the bloodstream.
8:06So pharmaceutical companies just had to design a drug that mimics that broken gene, right? A drug that blocks PCSK9. And that is exactly what happened. We now have FDA approved PCS canine inhibitors that are absolute blockbuster drugs for people who suffer from severe cardiovascular disease.
8:23That is incredible. And there are other examples, right? Many. We saw a similar pathway with a TYK2 gene, which was prioritized as a target for autoimmune diseases and led to a new oral medication for psoriasis.
8:35Wow. And of course, the BCL 11 aging. Genetic studies showed it suppresses fetal hemoglobin. which led directly to the development of Caskevy. Wait, Cass, Gabby, that's the... The world's 1st approved medicine based on Crisper Jean editing for sickle cell anemia, yes.
8:50That's huge. But this multi-scale genomics approach isn't just about ensuring a drug works, is it? It's also about predicting whether a drug is going to be toxic to you before you ever even swallow the pill.
9:01Exactly. The field of pharmacogenomics is becoming remarkably precise. Severe adverse drug reactions, which count for a massive portion of those clinical trial failures we talked about, are often driven by human leucocyte antigen or HLA variants.
9:15Right. Like if a patient carries the HLAB 57.01 variant, giving them a common antibiotic like flu clocksicillin can trigger severe liver injury. Exactly. And knowing that genetic cork exists allows doctors to simply choose a different antibiotic and completely avoid a life-threatening reaction.
9:33Okay, so if we're talking about these genetic quirks and broken genes, I want to pivot to something from the sources that completely reframed my understanding of DNA. Oh, what's that? Well, I always assume that the genetic variants causing disease we're sitting inside the genes that actually build our proteins, but most of the lariants identified by GLWs don't sit in the coding regions at all.
9:53They sit in the non-coding regions, like the vast stretches of DNA we used to dismiss as junk. Yes. Unlocking the non-coding genome has been one of the great aha moments of the last decade. We knew these non-coding variants were strongly associated with diseases, but because they didn't actively manufacture proteins, they were a complete mystery.
10:16But now we know what they do. We do. By integrating the single cell data we discussed earlier, we can finally see what they're doing. They are regulatory elements. It makes me think of like a movie set.
10:26Okay, how so? Well, for decades, we were so obsessed with the coding genes, right? The actors reading the script on camera. that we completely ignored the non-coding DNA. But those non-coding regions are actually the directors.
10:40That's a great way to look at it. They are standing off camera telling the actors to be louder, to be quieter, or to stop talking entirely. And the single cell data finally lets us hear the director's instructions.
10:50But we are learning they only give those instructions in very specific cell types. The director analogy captures it perfectly, especially when we look at complex conditions like obesity, because for a long time, the assumption was that the genetics of obesity were tied to adipose tissue.
11:05Meaning that fat cells themselves were genetically predisposed to store more energy. Right. But the data tells a completely different story. When researchers took a GWA for body mass index and mapped it against this new spadiocellular map, they found that the genetic variants driving obesity aren't enriched in fat cells at all.
11:24Where are they? They are heavily enriched in mid hypophylamic neurons. In the brain. in the brain. The implications of that discovery are profound. It proves that genetically obesity is fundamentally a neurological and behavioral condition.
11:38It is an energy balance equation dictated by brain signaling. The brain is telling the body, it is starving when it actually isn't. Exactly. Which entirely shifts how we view the condition. And incidentally, it completely validates the mechanism behind drugs like Ozempic and Wegovi.
11:54Absolutely, it does. Because those GLP1 agonists don't burn fat directly. They act on the receptors in the brain to regulate satiety. The genetic map finally proved why targeting the brain is the most effective way to treat weight loss.
12:08And we are seeing similar paradigm shifts in neurodegenerative diseases too. Like Alzheimer's. Yes. For years, Alzheimer's was viewed primarily through the lens of neuronal degradation, just brain cells breaking down, but multi-scale genomics revealed that Alzheimer's risk is actually heavily mediated by microglia.
12:26Microglia. Yeah, which are the primary immune cells of the brain. They act like the brain's trash collectors. So if the genetic directors in the microglia give the wrong instructions, the trash collector strike, and those beta amyloid plaques just build up.
12:39Exactly. It reframes Alzheimer's from a purely structural breakdown to an immune inflammatory failure. So, okay, if these non-coding directors are altering how single cells like neurons or microglia behave, how does that collective cellular chaos physically reshape an entire organ?
12:57Well, that is where the imaging data from the UK biobank becomes so critical. Researchers can now link specific genetic variants directly to macroscopic changes in the body. Like what kind of changes? Like, they can track the genetic low sci, responsible for thickening the regional walls of the heart's ventricles.
13:13In the brain, they can map the specific variants that accelerate biological aging or drive the physical accumulation of those amyloid plaques we just mentioned. This raises a massive question for me, actually, because I was looking at the findings on multi-organ connections in the sources.
13:30Yeah, that's a wild area of research. It is. The researchers discovered that the fractal dimensions of cardiac trebeculae, which is essentially the complex, spongy scaffolding on the inside of the heart muscle.
13:42They share the exact same genetic low sci with the dendritic complexity in the brain. If I'm a patient listening to this, and I realize my heart's physical structure, shares a genetic blueprint with my brain's physical structure, I'm suddenly wondering if a therapy designed to remodel my failing heart might accidentally alter the wiring in my brain.
14:02It is a critical concern, and it strikes at the heart of what biologists call pleiotropy. Yeah, pleotropy occurs when a single gene influences multiple seemingly unrelated physical traits. The challenge for researchers is disentangling that shared blueprint from an actual causal interaction.
14:22Like, are the heart and brain just built using the same genetic toolkit, or does a struggling heart actually cause the brain to degrade? How do you even begin to separate those 2 things? You can't just run a clinical trial where you intentionally induce heart failure to see what happens to the brain?
14:37No, you definitely cannot do it clinically, but you can do it genetically using a technique called Mendelian randomization. Okay. does that work? It essentially uses your genetics as a natural, randomized controlled trial.
14:49Because your genes are signed randomly at conception, researchers can group people by genetic risk, completely bypassing confounding lifestyle variables. Oh, that's clever. Very. And what they're finding is it is often a combination of both.
15:02There is playotropy, but adverse cardiac remodeling also causally accelerates neurodegenerative conditions. Wow, okay, so if they've isolated the malfunctioning heart muscle cell, I'm assuming the next step isn't just throwing drugs at it blindly.
15:18They must be looking for the specific gears inside that cell that are broken. Exactly. Let's walk through the case study on heart failure to see how all these scales integrate. Sure. So the multi-ancestry GWS for heart failure gave them the initial genetic coordinates.
15:33Then, layering on the single nucleus RNA sequencing revealed that the risk was enriched specifically in cardiomyocytes. Which are the heart muscle cells. Yes. So once they knew the exact cell, they moved down to the molecular scale.
15:46They conducted a proteom wide analysis connected to those cellular signals and identified 9 specific circulating proteins in the blood that are causally linked to heart failure. Nine proteins. That went from a vague statistical association across 1000000s of people down to 9 specific molecules.
16:03Incredible, isn't it? And those 9 proteins instantly become prioritized, genetically validated targets for a new heart failure drug. And looking at where this field is heading next, it feels like we are crossing into science fiction.
16:15The paper discusses an explosion of AI foundation models similar to large language models, but built for biology. Right, tools like SCGPT for single cell data or prot ST for protein sequences. Yeah. And how does that change things?
16:30The integration of these AI models with multi-scale genomics is driving just a breathtaking technological leap. We are moving toward a reality where researchers utilize techniques like perturb sec. Perturb sec.
16:43The sources mentioned that combines large scale CRISPR genetic screening with single cell transcriptomics. Wait, let me make sure I understand perturb sick. You're saying we can use CRISPR as a giant biological stress test.
16:56We can intentionally break or mutate 1000s of genes at once in a Petri dish and then use AI to watch exactly how the cell panics in response. That is the perfect way to describe it. You were intentionally stressing the system to map its ultimate breaking points.
17:10And doing this at scale is leading to the development of virtual cells. Yeah, by feeding all this multimodal data into foundation models, we can simulate human biology on a computer. We will soon be able to predict how a specific cell will respond to a novel drug in silico.
17:27In silico. So, before a single chemical is synthesized in a lab, let alone tested in a human patient. Exactly. As incredible as predicting drug responses in a virtual cell sounds, we need to ground this.
17:40Every major scientific leap has a catch. What are the limitations holding this back? The limitations are significant. First is the black box nature of these advanced AI models. Meaning we don't know how they get their answer.
17:53Right. One algorithm extracts a complex feature from a heart scan and predicts heart failure, it cannot always explain what that feature is in human biological terms. If a cardiologist cannot understand the physiological reason behind the AI's prediction, well, it is incredibly difficult to justify using it in clinical practice.
18:11That makes sense. And beyond the algorithmic black box, there is the fundamental issue of the genetic data itself. The sources made it very clear that our genomic databases are overwhelmingly skewed. It is arguably the most critical bottleneck right now.
18:23To perform the fine mapping required to identify true causal genetic variants. Researchers desperately need to expand ancestral diversity in these massive studies. Especially African populations, right?
18:35Yes, currently, there is a severe lack of data from African populations in particular. Human genetics is incredibly complex and diverse. If the foundational databases we use to train these virtual cells do not reflect global genetic diversity, the precision medicines we develop from them will not either.
18:54It is a mathematical and biological necessity to fix this and balance. So synthesizing everything we've covered today. Genomics is no longer just about handing a patient a vague disease risk score based on a statistical map.
19:05It has evolved into a multidimensional blueprint that connects molecular pathways to specific cellular behaviors, and ultimately to the physical physiology of our organs. By translating these genetic signals across all of these scales, we are fundamentally de-risking the entire drug development pipeline and accelerating the arrival of true precision therapies.
19:23What does this mean for the future of medicine? When your specific genetic blueprint dictates not just your risk of illness, but the precise cellular mechanism and physical organ structure, a new drug will target.
19:36It is a complete paradigm shift for how we treat human disease. And before we wrap up, there is one final concept from the research that I haven't been able to stop thinking about. The researchers discuss the idea of genetic aging clocks across different organs.
19:50Oh that's fascinating. Right. It raises a crazy question. Could your unique genetic blueprint dictate that your heart is biologically aging 10 years faster than your brain? And if multi-scale genomics maps those clocks perfectly, Could we eventually prescribe targeted drugs that reverse the biological age of one specific organ at a time?
20:10It fundamentally alters how we think about the aging process itself? It really does. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
20:23If you enjoyed this, follow, or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
20:38Thanks for listening and join us next time as we explore more science base by base.