Meta-analysis of MVP, UK Biobank and FinnGen with Mendelian randomization using eQTL/pQTL instruments implicates 6,447 genes and 69,669 causal gene-trait links.
0:00Welcome to Base by Base, the paper cast that brings Genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. I am really looking forward to getting into this one today.
0:10So, today, I want to start with a number that should honestly terrify anyone who is waiting for a new medicine. 90%. That is a really rough statistic. It really is. I mean, that is the failure rate of drug candidates entering clinical trials.
0:25Just imagine if an architect built 10 skyscrapers, you know, fully funded the construction, employed 1000s of people, and then watched 9 of them just collapse into a pile of rubble before a single tenet moved in.
0:38It's a staggering inefficiency. And the tragedy isn't just the money lost. You know, it's the time loss for the patients waiting for those drugs. Exactly. And the root of the problem isn't that we lack ideas.
0:48I mean, we have massive libraries of genetic data, 1000000s of variations linked to diseases. We have become very good at finding the scene of the crime in the genome. But we don't have the index. We really don't.
0:59We know a genetic variation is present at the crime scene, but we don't know if it's the mastermind, an accomplice, or just some innocent bystander walking their dog. Which is a crucial distinction. Because if you target the bystander with a drug, nothing happens to the actual disease, the building collapses.
1:17So here's the premise for today's deep dive. What if we could run a clinical trial without dosing a single patient? What if we could use nature's own randomized experiments, spread across 1200000 people, to pinpoint the exact biological switches that control human health?
1:35And that brings us to the study we are looking at today. They tested over 30000000 potential connections to do exactly that. 30 million. Yeah, it's a massive undertaking. And they didn't just find new targets.
1:47They actually managed to rediscover drugs we already use, but for diseases, we never expected them to treat. This isn't just some small academic exercise we're talking about, is it? Definitely not. Today, we celebrate the massive collaborative work led by Brian R.
2:00Faralito and his colleagues. This is a heavy hitter lineup. It really is. We are talking about a major effort involving the 1000000 veteran program, or MVP, the VA healthcare system, the broad institute of MIT in Harvard, plus collaborators from the UK BioBank and Finn Gan.
2:19Basically, the Avengers of Biobanking. Pretty much. They are integrating data from the largest biobanks on the planet, and that scale is entirely non-negotiable here. Why is the scale so important? Because when you are trying to cut through the noise of human biology, to find a real causal signal, you need 100s of 1000s of participants, or the math just doesn't work out.
2:41Okay, let's set the stage on the problem they're actually trying to solve. For those of you listening, we've discussed GWS, do you know wide association studies on the show before? Right. We've gotten really efficient at scanning the genome and finding those dots that light up for certain diseases.
2:54So we know where to look generally. We know the neighborhood. We can say people with this genetic marker are statistically more likely to have high blood pressure. But that's the crucial gap, isn't it?
3:03Yeah, because GWA doesn't give you the mechanism. It tells you there is an association, but it doesn't tell you what to do about it. Because a genetic marker is just a signpost? Exactly. It doesn't tell you if the gene is cranking up a protein or shutting it down entirely.
3:17And if you're making a drug, you absolutely need to know. that You do? Do I block this pathway or do I boost it? If you guess wrong, you just spend a $1000000 to make a drug that does the exact opposite of what you wanted.
3:29Which is an incredibly expensive coin flip. It is. So to solve this, we have to integrate GWAs with Olmecs, specifically transcriptomics, which is the RNA or the recipe and proteomics, the actual proteins or the cake.
3:44Moving from association to causality. That is the whole mission of this deep dive. Harmonizing all that data to prioritize drug targets that have actual genetic validation. The statistic that jumped out at me from the reading was about the success rate when you actually have that validation.
4:00It's wild. Historically, if a drug target has genetic evidence backing it up. It is twice as likely to succeed in clinical trials compared to one that doesn't. Doubling your odds in a 90% failure industry is huge.
4:13It's a total game changer. So let's unpack the how. How do you physically do this without running a trial? I guess it starts with the data set size. The scale is the engine here. They meta-analyzed over 1200000 individuals across 2003 different phenotypes.
4:31Which is just scientific shorthand for traits or diseases. Right. And the methodology they used is Mendelian Randomization or MLER. I've heard this describe as nature's clinical trial, but that can sound a bit, you know, abstract.
4:43How does it actually work in practice? Well, think about a standard clinical trial? A researcher flips a coin to decide if you get the drug or the placebo. And that randomization is magic because it removes bias.
4:53Exactly. It ensures that the 2 groups are identical except for that one drug. It prevents things like, oh, the people taking the drug also happen to exercise more, from ruining your data. So you know that the result is because of the drug, not because one group was younger or ate healthier.
5:08Now, in Mendelian randomization, nature flips the coin for us at conception. When you inherit your DNA. Exactly. It's completely random. Some people inherit a variant that naturally raises the level of a specific protein in their blood just a little bit higher than average.
5:26Well others inherit a variant that lowers it. Precisely. Okay. So these people have essentially been on a natural dose of that protein their entire lives. Wow. Yeah. So if we look at the people who naturally have higher levels of protein acts because of their genetics, and we see they have a significantly lower risk of heart disease.
5:46We can infer causality. Yes. We can say protein X likely protects against heart disease. And it cuts right through the confusion of lifestyle factors because your genes don't care if you smoke or jog. They are fixed at birth.
5:58But they didn't just look at the genome and the disease. They looked at the steps in between. Yes, this is where those instruments come in. Yeah, can't just draw a straight line from gene to disease. You need the bridge.
6:07Right. So they looked at EQTLs, which are expression, quantitative trait lossi. That measures how much RNA is being produced. The instruction. Exactly. And they also looked at PQTLs, protein quantitative trait losi, which measures how much the protein is actually floating in the blood.
6:24They pulled this from massive databases too, didn't they? G-Tex, Eric, decode. Yes, those are basically the gold standards for this kind of molecular data. And the computational lift here must have been absurd.
6:35It was. They performed 2 sample MR on every single combination that is 31.5 million unique gene trade associations tested. 31.500000 tests. I mean, if you did that manually, you'd be done in about 3000 years.
6:50Easily. But you can't just trust every result a computer spits out, obviously. There's noise in the data. Tons of noise. That why they applied strict filtering. They look for concordance. Right. Imagine you have 3 witnesses to a crime.
7:02One is the RNA data. One is the protein data from the UK, and one is protein data from Iceland. And if one says the suspect is tall and the other says the suspect is short, you throw the case out. You only want the cases where everyone agrees on the description.
7:17They only kept signals where different data sources agreed on the direction. Like, does more protein equal more disease? If the data conflicted, they tossed it. Exactly. They needed robust concordant signals.
7:31Now, this is where it gets really cool for me. They didn't just stop at finding links. They basically built a drug hunter robot using machine learning. They used XG boost, which is a gradient boosting algorithm.
7:43But the tech itself isn't as important as their strategy. What do you mean by their strategy? They needed a truth set. So they went to the Chimble database and pulled a list of all approved successful drugs.
7:54So they gave the computer the answer key. In a way, yeah. They told the model, here is what a winner looks like genetically. These are the genetic patterns of targets that actually became drugs. Now look at our 31000000 new test results and rank them based on how much they resemble these known winners.
8:11It's like training a sniffer dog. You give it the scent of the target and send it into the woods. So what did the dog find out of those 31 million? They narrowed it down to 69,669 gene trait pairs with strong causal evidence.
8:25That's a massive reduction. It is, and the statistical bar was incredibly high. A P value less than one. 6 times 10 to the -9. Which is, well, practically zero. It means there is virtually 0 chance of it just being a coincidence.
8:39Okay, but I am going to play the skeptic here for a second. Anyone can generate a list of 70,000 things and say these are important. How do we know this method actually works in the real world? That is the validation step.
8:52And it's crucial. To prove the model works. They checked if it could find drugs we already have. It's called rediscovery. So if the model is smart, it should be able to look at the data and say, hey, this HMGCR gene looks like a great target for cholesterol.
9:07Which we know is true because that's what statens target. And it did exactly that. It rediscovered 9% of all approved drug targets purely through this blind data analysis. 9%. Initially, that sounds kind of low.
9:19I feel like if I built a robot to find cars and it only found 9% of the cars in the parking lot, I'd be pretty worried. But context is everything here. Remember, they are looking at the entire genome blindly.
9:30They aren't looking in a parking lot. They're looking at the whole city from space. Finding 9% of the entire pharmacopoeia without knowing what you are looking for is actually wildly impressive. That makes a lot of sense when you put it that way.
9:43Plus, look at where it succeeded. It was incredibly good at finding cardiovascular targets because we have excellent data on lipids and blood pressure. And less good at cancer. Right. Cancer is complex, it's localized.
9:57And the data sets just aren't as robust for those specific pathways yet. But there's a kicker regarding the mechanism, right? Oh, definitely. For the drugs it did rediscover. The genetics correctly predicted the mechanism of action, meaning whether it's an inhibitor or an activator, 84% of the time.
10:13That brings us back to that expensive coin flip. If this method tells you to block a protein, you can be 84% confident that blocking is the right move, not boosting. Exactly. That saves years of failed experiments.
10:25It provides directional certainty, which is just invaluable. Let's talk about the treasure hunt aspect of this. Repurposing. Finding old drugs that can learn new tricks. This seems like the absolute lowest hanging fruit for pharma.
10:38It is because these drugs are already proven safe in humans. The study found 3364 potential repurposing opportunities. Give us the highlights. What really stood out to you? Well, take Matt Foreman. It's the frontline drug for type 2 diabetes.
10:54We consume tons of it globally. Yeah. The data suggests it has a causal link to reducing atrial fibrillation. Hey, fib, that's a heart rhythm issue. That feels pretty disconnected from blood sugar. You would think so.
11:04But it suggests there's a metabolic component to heart rhythm that we are underestimating. Fascinating. Then there is Coastal Zoomab. That's an arthritis drug. It targets the I 6 receptor to lower inflammation and joints.
11:16And the model flag that for AFIB too. Yes, it did. So we are seeing inflammation pathways popping up in heart conditions. Precisely. blurring the lines between our medical specialties. Inflammation is inflammation, whether it's in your knee or your heart atria.
11:29The genetics don't care about our medical textbook chapters. Exactly. Another really interesting one was Enru Kinzumab, which is an IL 13 antagonist usually looked at for colitis. The data flagged it as a strong candidate for psoriasis.
11:43It really emphasizes how connected these immune systems are. It does. Speaking of the heart, the paper had a specific vignette or a case study on lipids that seem to be the proof of concept for their whole approach.
11:55Yes, the dyslypodemia dive. They wanted to see if they could find a brand new way to lower cholesterol using this method. And they recovered the known hits like PCSK9 and HMGCR. Right. We know PCSK9 is the bad guy.
12:09It destroys the receptors that clear bad cholesterol from your blood. we already have drugs that inhibit it. Okay, so that's the control. I found the stuff we know, but what was the new find? They identified a high probability target called ANXA2, or NXA2?
12:23This isn't a drug target we currently use. What does ANXA 2 actually do? The mechanism is really elegant. NXA2 is a natural inhibitor of PCS K9. Wait, let me get this straight. So PCSK9 inhibits the cholesterol cleaners.
12:36Yes, it destroys them. And ANXA 2 inhibits PCSK9. Correct. It's double negative. NXA 2 inhibits the inhibitor. So it's basically taking the brakes off the cleaners. That is a perfect way to look at it.
12:49So logically, if you can boost ANXA 2 or mimic its effects, you stop PCSK9 from doing its damage. Which lowers the cholesterol. Exactly. And the model flagged this as a top tier candidate. It's effectively handing pharma a roadmap for a new class of cholesterol drugs.
13:07That's the aha moment. Nature essentially already has a drug for high cholesterol inside us. We just need to figure out how to bottle it. It validates that the method can find biological logic, not just random statistical spikes.
13:19So zooming out a bit. If I'm sitting in a boardroom at Pfizer or Merk right now, Why does this paper matter to my bottom line? It matters because it shifts your betting odds. Instead of guessing based on a mouse model that might not translate to humans at all, you are prioritizing targets that have human genetic validation.
13:36You are placing smarter bets. Exactly. It's an efficiency roadmap. But it's not just about finding hits, is it? It's also about avoiding disasters. That's the other side of the coin. safety. The data reveals risks just as clearly as benefits.
13:51For example, the study flagged the target of a drug called Trastuzamab. That's a breast cancer drug, right? Extremely effective for HER2 positive cancer. Yes, but clinically, we know it has a nasty side effect.
14:04Cardio toxicity. It can cause heart failure. And the genetic method picked this up blindly. It did. The algorithm flagged the gene associated with Trastazumab as being causally linked to heart failure.
14:15That is incredible. It implies that if we had run this analysis before the drug was ever invented, we would have known about the heart failure risk. Exactly. It allows us to anticipate toxicity. If you see a genetic link to a serious side effect.
14:28You design your clinical trial differently. You monitor hearts from day one. Or if the risk is simply too high, you kill the program before you spend a billion dollars. It turns unforeseen side effects into foreseen risks.
14:41We do have to be realistic though. This sounds like a crystal ball, but it can't be perfect. Where does it fail? Oh, it is absolutely not perfect. We mentioned directionality earlier. Getting that block versus boost decision, right, is tricky.
14:54even with that 84% success rate. But the biggest limitation is tissue specificity. Because most of this data comes from blood draws. Exactly. The PQTL data, the protein levels are mostly from plasma. But biology is local.
15:08What do you mean by local? Well, a protein might be doing one thing in your blood but something totally different inside your brain or your liver? Got it. So if we are looking for an Alzheimer's drug, a blood protein level might not reflect what is actually happening behind the blood brain barrier.
15:24So assuming blood equals body is a dangerous simplification. It is. And that's why the authors really emphasize triangulation. They didn't just rely on this one MR method. What else did they use? They combine the results with databases like OMIM, which tracks rare Mendelian genetic diseases and mouse knockout databases.
15:42Mouse knockouts are where they delete a gene and a mouse to see what breaks, right? Yes. The study found that if a gene hits in their MR analysis and shows up in a mouse model. The odds of it being a valid drug target skyrocket.
15:56Like building a legal case. Exactly. DNA evidence the MR is great. But DNA evidence plus a fingerprint, the mouse data, plus an eyewitness, the omim data, is unbeatable. It moves genomics from a descriptive science, you know, just cataloguing a list of genes to a predictive tool.
16:16It finally answers, what can we actually do with this? That is the core shift here? And it leaves us with a really provocative thought. Think about the 1000s of failed compounds sitting in pharmaceutical libraries right now.
16:28The dusty shelves of the valley of death. Exactly. Compounds that were safe, but failed efficacy. They simply didn't work for the disease they were tested on. But maybe they were just tested on the wrong disease.
16:38Precisely. This method suggests that the failure wasn't necessarily the molecule. It was the question we asked it. Wow. Maybe that failed asthma drug is actually a breakthrough for heart disease, and we just never looked at the right genetic map to tell us so.
16:50That is a thought to keep you up at night in a good way. The cures for diseases we consider incurable might already be sitting in a freezer somewhere, just waiting for the right data to unlock them. What a fascinating place to leave it.
17:02It really is an exciting time for genomics. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
17:14If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode, and inspired by the article you've just heard about.
17:29Thanks for listening and join us next time as we explore more science base by base.