Zuber et al. introduce MrDAG, a Bayesian causal graphical model that combines Mendelian randomization, structure learning, and interventional calculus to estimate causal effects among multiple correlated exposures and outcomes using summary-level GWAS data. The method reveals dependency structures and highlights education and smoking as key intervention points for mental health.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So imagine discovering that the reason a patient developed a really severe mental health condition, like schizophrenia.
0:14Had less to do with their direct lifestyle choices, and almost entirely to do with a hidden genetic traffic jam happening in a completely different part of their biological network. Right, which is, it's a wild thing to think about.
0:27It really is. I mean, we look at the world around us, and we observe this staggering scale of the mental health crisis, where one in 8 people globally suffer from some form of mental health phenotype. And that accounts for over 15% of total years lived with disability.
0:42So naturally, we try to draw straight lines. Yeah, we crave simplicity. We assume a lack of sleep directly causes depression or that, you know, smoking directly triggers psychotic disorders, but human biology, well, it actively resists those straight lines.
0:57totally defies them. Exactly. When we look at the genetics of mental health and lifestyle behaviors, we're looking at this deeply interconnected system where symptoms often masquerade as the root causes.
1:08Right. Right. So here is the question for you, the listener. What really happens when you try to untangle this massive web of overlapping lifestyle habits and mental health conditions using pure genomic data?
1:21Like, how could this change our understanding of what is actually causing a disease versus what is just a biological side effect? Because that's the whole point of this deep dive, right? We're exploring a new analytical framework that fundamentally rewrites the map of mental health liabilities.
1:37Yeah, a methodology that separates the true biological culprits from the innocent bystanders caught in the crossfire of our genetic code. So today we celebrate the work of Arena Zuber, 20 Cronje, Nakai, Depender Gill, and Leonardo Botolo, from institutions, including Imperial College, London, and the University of Cambridge, who have advanced our understanding of complex pathways toward disease using joint Medelian randomization.
2:00It's a fantastic paper. It is. And this research was published in the American Journal of Human Genetics, volume 112, on May 1st, 2025. So to really grasp the leap forward this research represents, we have to establish the wall that genetic epidemiology has been hitting for the last decade.
2:19Right, the wall. Yeah. Historically, to figure out if an exposure, say, smoking causes an outcome like depression, you observe a population. But those observational studies are, well, they're perpetually haunted by confounding variables.
2:34The classic correlation versus causation tract. You observe that people who smoke have higher rates of depression. But you don't know if the smoking biologically causes the depression, or if maybe people with depression are using nicotine to self-medicate.
2:48Or if there's a 3rd completely unmeasured factor. Right, right. Like chronic socioeconomic stress that independently causes both the smoking and the depression. Yeah, that 3rd unmeasured factor is the confounder, and it is the absolute bane of observational science, which is why the field turned to Mendelian randomization, or MR attempts to bypass those environmental confounders by using a person's genetic variations as a proxy for the exposure.
3:16It uses genetic markers as instrumental variables. Okay, let's unpack this with an analogy to make sure we're grounded here. Sure. Using a genetic instrumental variable is kind of like a randomly assigned lottery ticket given to you at birth.
3:29I like that. Right. So if we want to study the effect of smoking on depression. Instead of looking at whether you actually smoke, which is influenced by your stressful job or your peer group. We look at whether you drew the genetic lottery ticket that predisposes you to nicotine addiction.
3:45Because your genes were locked in at conception. Exactly. They can't be altered by your stressful job later in life. That is the core logic of the MR paradigm. By relying only on the genetic predisposition to an exposure, the environmental noise just, it falls away.
4:00We get to look at the pure biological blueprint. sounds great. It is, but traditional MR has a really severe limitation. It generally operates on a binary A to B framework. It tests one exposure against one outcome at a time.
4:15And as we just said, human viability is not a series of isolated one-way streets. No, it's trying to understand mental health with traditional MR is like trying to fix a city's traffic problem by only staring at one single intersection.
4:31Oh, that's a great way to put it. You might think you found the source of the jam, completely missing that a bridge is closed 3 miles away. In genetics, this interconnectedness creates a massive problem.
4:41A huge problem. We call it pleotropy. The genetic variants we use as our lottery tickets, well, if they often do more than one job in the body. A gene that predisposes you to smoke might also independently affect your sleep cycle through a completely separate biological pathway.
4:56Which money is the waters? Exactly. So when scientists realize this, they developed multivariable Mendelian randomization, MVMR. To analyze multiple exposures at once. Right. But even MVMR hits a wall when the exposures start interacting with each other.
5:10If education levels influence your likelihood to smoke and smoking influences your sleep, MDMR really struggles to figure out the hierarchy of those dependencies, it just sees a cluster of correlated genetic signals.
5:22Which brings us to the methodological breakthrough of this paper. The researchers developed a Besian causal graphical model called Mitterdag to map the entire grid simultaneously. Yes, Mr. Dag. Let's break down that acronym for the listener.
5:36The day G stands for directed a cyclic graph. In data science, a graph isn't like a bar chart. It's a visual web of nodes connected by lines. Directed means those lines have arrows indicating a one-way street of cost and effect.
5:50And a cyclic means there are no impossible time traveling loops. You can't have exposure, A, cause B, B, cause C, and C turnaround and cause A, the biological timeline has to flow forward. Perfectly explained.
6:03And the mentrodeg algorithm operates on 3 interlinked strategies or pillars to build this graph. First, it anchors itself in the MNR paradigm we just talked about. It uses summary level genetic data from large biobanks as its instrumental variables.
6:17This enforces a strict directional boundary. Genetically predicted lifestyle exposures can cause mental health outcomes, but the mental health outcomes can't reach back in time and alter your genetics.
6:30Okay, so that establishes the foundation. We're dealing with pure genetic signals. immune to environmental noise. But that still doesn't tell us how a factor like education and a factor like smoking connect to each other within the network.
6:44How does the model figure out which way the traffic is flowing between the exposures themselves? That requires the 2nd pillar, structure learning. Structural learning. Yeah. Instead of human researchers guessing the connections, the algorithm mathematically interrogates the data to learn the web of dependencies, it calculates the conditional probabilities across all the variables simultaneously.
7:06Yeah, it figures out how the different exposures interact and crucially how the genetic liabilities for different mental health outcomes interact with one another. But wait, I need to push back on this.
7:15How can a computer algorithm definitively know what's a cause and what's just a really tight coincidence without running a real world physical trial? It sounds like we're trusting a math equation to solve the physical reality of human disease.
7:29I get that. But what's fascinating here is how the algorithm achieves this through the 3rd pillar. Interventional calculus. Interventional calculus. Right. It's based on computer scientist Judea pearls do calculus.
7:42There is a fundamental mathematical difference between seeing a correlation and doing an intervention. Okay, give me an example of the difference between seeing and doing. Well, think of a barometer in a storm.
7:53If you see the needle on a barometer drop, you can confidently predict that a storm is coming. They are highly correlated. Right, obviously. But if you mathematically do an intervention, like if you manually reach out, grab the needle of the barometer and force it to drop, you don't cause a storm to materialize in the sky.
8:11Ah, because the dropping needle is just a downstream symptom of the atmospheric pressure, which is the true root cause. Exactly. Metrodag's algorithm performs mathematical interventions on the genetic data.
8:23Once it learns the structure of the network. It simulates what would happen if it forced one specific node lake smoking to change while holding the rest of the web constant. If it forces the smoking needle to move, and the depression outcome moves with it, it identifies a true causal pathway.
8:42If it moves the needle and nothing else happens, it knows it's just looking at a barometer. Okay, that distinction between seeing and doing mathematically makes so much sense. But obviously the researchers didn't just write this algorithm and immediately assume it worked perfectly on human biology, right?
8:58No, they subjected Meptodag to some severe stress tests. They tested it against the most advanced models currently used in the field, specifically MR, BMA, and a network MR tool called MR2. And to do this, they created highly complex synthetic data sets to see if the algorithms could be tricked.
9:15And the most punishing of these simulations was called the DagXDGY scenario. Yes. In that scenario, the researchers simulated a deeply entangled network. They created a web of primary exposures that cause secondary exposures, which, in turn, caused multiple outcomes, which then caused secondary outcomes.
9:34Just a labyrinth of pleotropy. Total labyrinth. And when they ran the older models. on the synthetic data, those models through false positives everywhere. They saw 2 things moving together at the end of the chain and mistakenly declared a direct causal link between them.
9:50Because the older models couldn't see the hidden hubs mediating the traffic, they were just looking at the final destination and drawing a straight line back to the start. Precisely. But Mr. Dagg, because it models the dependencies within the exposures and within the outcome simultaneously, recognize the mediators.
10:06It successfully shrank the false positive causal relationships to zero. Wow. 0 false positives. Yeah. It didn't get fooled by the biological barometers. That's incredible. And once they proved the math worked in these chaotic simulations, they brought Mr. Dag into the real world.
10:21They pointed this tool at the immense complexity of psychiatric genetics. Yeah, they utilized massive summary level GWAS data genome wide association studies covering 100s of 1000s of individuals. They fed the algorithm data for 6 major lifestyle exposures.
10:38Okay, let's list those. Genetically predicted education, physical activity, sleep duration, alcohol consumption, smoking, and leisure screen time. Right. And those are the 6 pillars of modern public health advice.
10:51We're constantly told that manipulating those 6 things will dictate our health. Exactly. So they tested those 6 against 7 major mental health phenotypes, major depressive disorder, ADHD, anorexiandervosa, bipolar disorder, autism spectrum disorder, schizophrenia, and a measure of general cognitive performance.
11:09So we have a grid of 6 exposures and 7 outcomes. Right. And older single intersection models would look at this data and draw lines everywhere. They'd say sleep effects, depression, screens affect ADHD, alcohol effects, schizophrenia.
11:22But when Mr. Dagg applied its structure learning and interventional calculus to filter out the noise, the resulting map was astonishingly sparse. Out of all those complex lifestyle factors, Mr. Dagg revealed that only 2 emerged as significant root cause points of intervention.
11:40Education and smoking. Just those two. But wait, only education is smoking. What about leisure screen time? I mean, the prevailing narrative right now is that screen time is a primary driver of the mental health crisis.
11:51I know, and this is where the power of the DAG network really becomes visible. Mr. Dagg showed that factors like screen time and sleep duration are not root biological causes for these genetic liabilities.
12:02No, they are downstream effects, mediated by other factors in the web. For instance, the data revealed that a higher genetic liability for education has a negative causal effect on both smoking and leisure screen time.
12:14So if you only look at the intersection of screen time and mental health, you assume the screens are the cause. But Mr. Dagg zooms out and realizes the screens are just a symptom. The true causal momentum is originating further up the chain at the education node.
12:29It completely forces us to reconsider where we intervene. If you want to change the downstream outcomes, you have to target the upstream hubs. Okay, here's where it gets really interesting. Let's focus on a specific pathway that perfectly illustrates this, because this finding regarding schizophrenia is a massive paradigm shift.
12:48really is. For years, observational epidemiology, and even some earlier genetic studies, claimed there was a direct biological causal link between smoking and the development of schizophrenia. The data always showed a thick, bold line connecting the two.
13:04Yeah, that correlation has been a longstanding fixture in the literature. But when Mr. Dagg mapped the dependencies within the mental health outcomes themselves. It uncovered a hidden path that completely dismantled that old narrative.
13:16Okay, walk us through the mechanics of that. How did Mr. Dagg prove the straight line was an illusion? Well, when the algorithm mapped the network, it found that the genetic liability for smoking does not point directly to the genetic liability for schizophrenia.
13:30It doesn't. No. Instead, smoking directly increases the genetic liability for major depressive disorder or MDD. Okay, so the 1st link in the chain is smoking to MDD. Right. Then MDD acts as a massive central hub in the psychiatric network.
13:45The genetic liability for MDD directly increases the liability for bipolar disorder. And finally, bipolar disorder directly increases the liability for schizophrenia. So it's a chain reaction. Smoking to MDD to bipolar to schizophrenia.
14:00But how does the math prove that the direct smoking dischizophrenia link is false? This brings us back to the interventional calculus. The algorithm mathematically holds the MDD node constant. It basically says, let's freeze MDD so it can't fluctuate, and then let's simulate a change in smoking.
14:17When it does that, the genetic correlation between smoking and schizophrenia drops to absolutely zero, the signal vanishes. Wow. That proves the causal signal was never a direct line. It was traveling through MDD the entire time.
14:28That fundamentally alters how we view the disease architecture. The smoking was causing a genetic traffic jam at the MDD intersection, which eventually backed up into bipolar, which then spilled over into schizophrenia.
14:40Exactly. And the older models just saw a car entering the highway at smoking and exiting its schizophrenia, and assumed they drove straight there. Yeah, the clinical implications of this are just immense.
14:51If we connect this to the bigger picture. Enderdeg identified major depressive disorder as a superhub. It's a central mediating node for multiple psychiatric phenotypes, not just the pathway to schizophrenia.
15:05So practically speaking, what does this mean for preventative psychiatric medicine? It suggests that if clinical interventions, whether pharmacological or behavioral could effectively treat or prevent MDD, you wouldn't just be curing depression.
15:18Right, because of the cascade. Exactly. Because MDD mediates so many of these pathways, intervening at that central hub could theoretically halt the downstream cascade, preventing the subsequent development of bipolar disorder, schizophrenia, or even anorexia nervosa, you cut off the fuel supply to the rest of the network.
15:36That makes such a compelling case for aggressively prioritizing early intervention for depression, not just as a standalone condition, but as the gateway to the broader psychiatric network. Right. But, you know, as we analyze this map, I have to point out a finding that seems highly counterintuitive.
15:54The data output says that higher genetically predicted education directly increases the risk for autism spectrum disorder or ASD, from a biological standpoint, does going to school actually cause autism?
16:08That doesn't make intuitive sense. No, it doesn't. And this is a phenomenal point to raise. It highlights why algorithmic output still require deep domain expertise to interpret. We have to look at the data source.
16:18Mentor Dag is analyzing summary statistics from genome wide association studies, drawing from massive databases like the UK Biobank. Right, these huge repositories of genetic and medical data. But think about how someone gets flagged as having ASD in those databases.
16:32The GWS data relies on clinical diagnostic codes. This introduces a well-known issue in genomics called ascertainment bias. Ah, okay. As for damn it, bias. So it's about who is actually getting counted.
16:43Yes. The genetic markers predicting higher educational attainment do not cause the underlying biological neurodivergence of autism. However, an individual who is embedded in a standardized formal educational system for a longer period is subjected to far more institutional observation.
17:02Oh I see. They are significantly more likely to have their neurodivergent traits recognized, formally evaluated, and officially diagnosed by medical professionals. So the algorithm is technically correct based on the data it was fed.
17:15It sees a strong causal link between education and the presence of an ASD diagnostic code. Exactly. But as researchers, we have to synthesize that and recognize it's a sociological pathway, a diagnostic bias, not a biological mutation caused by reading textbooks.
17:30Right. The typical age of onset for core ASD traits precedes advanced formal education. So biologically, the timeline doesn't support education as the origin. The structured environment of the education system serves as the mechanism for diagnostic capture.
17:45That makes total sense. Tools like Mr. Dag are incredibly powerful for mapping the network, but we must apply our understanding of ascertainment bias and biobank architecture to interpret why the arrows point the way they do.
17:57It's a vital reminder for anyone working with GWO's data. The map is only as accurate as the territory it's drawn from. If our healthcare system has biases and who gets diagnosed, those biases will be digitized, embedded in the biobank, and eventually output by the algorithm as a causal link.
18:16And the authors of this paper are highly transparent about this limitation. It really highlights the need for future iterations of precision medicine to better distinguish between the core biological traits of a condition and the socioeconomic likelihood of receiving a formal diagnosis.
18:30As the underlying biobank data becomes more refined, the causal maps generated by Mr. Dagg will become even more precise. So bringing this all together, what we are looking at is a massive leap forward in how we model human health.
18:42By utilizing Beesian causal graphical models and relying purely on genetic instrumental variables. MintterDag allows us to bypass the noise of environmental confounders. really does. It proves that we can no longer isolate single variables in a vacuum.
18:56By mathematically mapping the hidden connections between our behaviors and our genetics, it reveals that some causes of disease are actually just downstream symptoms of something else entirely. It transitions genetic epidemiology from a flat two-dimensional view of isolated correlations into a rich three-dimensional reality of biological cascades.
19:18So what does this all mean for you? The next time you read a flashy headline claiming X causes Y? Ask yourself, are they looking at the whole traffic grid or just one intersection? What other hidden biological hubs might be steering your health without you even knowing?
19:33This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
19:47If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
19:55Thanks for listening and join us next time as we explore more science, base by base.