A 38-cohort proteogenomic meta-analysis of up to 78,664 people maps fine‑mapped protein quantitative trait loci (pQTLs) across 1,116 circulating proteins, uses machine learning to assign trans effector genes, and triangulates genetic and observational evidence to highlight disease mechanisms and therapeutic opportunities.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine you're uh, managing the architecture for a massive global cloud computing network.
0:15Right, like a sprawling AWS setup or something. Exactly. And suddenly, an entire cluster of user facing applications in Tokyo just completely crashes. So naturally, your 1st instinct is to inspect the local hardware in Tokyo.
0:29Yeah, you check the physical servers, the local code, the regional power supply. But everything looks perfectly fine. Which is always the most frustrating part. Oh, absolutely. So after weeks of tearing your hair out, you finally discover the real culprit.
0:41And it's a subtle misconfiguration in a global routing protocol. Ah, let me guess. Pushed from a server farm, 1000s of miles away. You got it. From the server in Virginia. So the failure was local, but the cause was systemic and incredibly distant.
0:56That is a perfect analogy for what we're looking at today. Right. Today, in this deep dive, we are exploring a complete paradigm shift in Genomics that mirrors this exact scenario. We're talking about the protein circulating in your blood.
1:09The very ones that determine whether you develop complex diseases. And what we are finding out. Is it the vast majority of their genetic regulation isn't happening locally at all? No, it's really not. The breakdown is being orchestrated by distant, hidden genetic switches.
1:25So the big question for you, the listener, is this. How could mapping this sprawling, decentralized network completely change the way we validate biomarkers, repurpose drugs, and treat diseases like rheumatoid arthritis, or heart failure?
1:41Well, to map a network of that magnitude, we are examining a truly unprecedented multi-cohort analysis. Yeah, today we celebrate the work of Mind Copperloo, Carl Smith Byrne, Brian Richard Faralito, Claudia Langenberg, Make Pietzner, and a massive international consortium.
1:57It's an incredible team. Researchers from 38 different cohorts who have advanced our understanding of the human prodium. And specifically it's genetic regulation. Their 2026 paper, published in the journal cell, provides the foundation for our deep dive today.
2:11It is, um, it's crucial to anchor a conversation on why this massive consortium targeted circulating proteins in the 1st place. Right, because we already have 1000s of genome wide association studies, don't we?
2:23We do. And they highlight regions of our DNA linked to disease risk, but DNA is static. It's just the architectural blueprint. Exactly. Proteins are the dynamic physical infrastructure of the body. They're the structural beams, the signaling molecules, the enzymes.
2:39So the gap in our knowledge has basically always been tracing that path. Like from a genetic variant in a non-coding region of DNA, all the way to the actual physical manifestation of a disease. Yeah, and proteins in the blood serve as that critical intermediary layer.
2:55Okay, let's unpack this. Because historically, we've had a very um, a very localized view of that intermediary layer. Wow, very localized. One geneticist searched for the cause of a malfunctioning protein, they almost exclusively hunted for cis variants.
3:09And just to clarify, cis variants are genetic mutations located right next door to the gene that actually encodes the target protein, like the worker standing right at the assembly line making the part.
3:20That's right. But that localized view has been incredibly limiting. It assumes biology operates on a simple, isolated linear script. Wait, isn't the body way more complex than that? Like, what about the managers in a completely different building who ordered the parts?
3:35Why have we been ignoring them? Well, we haven't been ignoring them on purpose. It's just really hard to map, but human biology is fundamentally a complex integrated system. Which brings us back to that cloud computing analogy.
3:47A cis variant is a broken server in Tokyo causing a crash in Tokyo. Exactly. But what we've been missing are the transvariants. The managers in the other building. Right. These are the distant genetic regulators sitting on entirely different tromosomes.
4:01The servers in Virginia manipulating the network in Tokyo. And I want to pause here because we need to demystify how a transvariant actually works. I mean, it isn't a magical remote control. There has to be a physical mechanism, right?
4:13Absolutely. The mechanisms are highly physical and incredibly diverse. How so? Well, often a transvariance sits within a gene that codes for a transcription factor that's a protein that physically travels through the cellular environment.
4:27Oh, so it actually moves. Yes, it enters the nucleus and binds to the promoter region of our target gene on another chromosome to dial its expression up or down. Wow. Okay, so it's a physical delivery system.
4:39Or alternatively, it might involve massive three-dimensional chromatin looping. Chromatin looping. Yeah, where the physical architecture of the DNA folds upon itself. It brings a distant enhancer region into direct physical contact with our target gene.
4:54That is wild, but capturing those systemic interactions requires just a staggering amount of statistical power. Oh, completely. If you are looking for connections between 1000000s of genetic variants and 1000s of circulating proteins across different chromosomes.
5:10The sheer volume of data creates a massive multiple testing problem, right? The noise threatens to drown out the signal. Which is why the scale and the strict methodology of this study are so vital. We are looking at data from up to 78,664 participants.
5:25Almost 80,000 people. huge. Across 38 independent studies. And they measured over 1,100 circulating proteins in the blood. And I think in the UK, biobank cohort alone, they measured up to 1463 distinct proteins.
5:39Yeah, they generated this proteomic data using OLink technology, which relies on antibody-based proximity extension assays. The mechanics of this are actually quite elegant. Instead of just using a single antibody that might, you know, non-specifically bind to the wrong target and body the data.
5:56Which happens a lot. Olink uses a dual antibody system. Two distinct antibodies, each tag with a unique short DNA sequence, must both bind to the exact same target protein. And that dual binding requirement is what provides the incredible specificity.
6:10Because of only one binds, nothing happens. Exactly. When both antibodies bind, their DNA tags are brought into close proximity. They hybridize creating a unique DNA barcode. That's so clever. And then that barcode is exponentially amplified using PCR.
6:24You got it. It basically translates a microscopic protein concentration into a massive, highly quantifiable genomic signal. Okay, so we have incredibly precise protein quantification across nearly 80,000 people.
6:39I mean, how did they avoid false positives with so much data? Well, the researchers utilized inverse variants, fixed effects meta analyses. That's how they located the protein quantitative trait lowci, or PQTL.
6:53And a fixed effects model is a deliberate choice here, right? It is. They are operating under the assumption that the underlying genetic effect on a proteins expression is fundamentally the same across these cohorts.
7:03Which, we should note, were predominantly of European ancestry. Yes, that's an important caveat. And the inverse variants waiting simply ensures that studies with larger sample sizes and more precise data have a stronger pull on the final statistical consensus.
7:17So identifying a PQTL, a coordinate on the genome correlated with a protein level is one thing, but they took it a step further. They had to. In the tangled web of a trans association, a single genetic signal might encompass dozens of genes.
7:30That's linkage to equilibrium, right? Exactly. To solve this, they deployed machine learning guided effector gene assignment. So they didn't just guess? No, they integrated chromatin state data, gene expression data, and algorithmic predictions to pinpoint the specific causal distant gene pulling the strings.
7:48And we really need to emphasize their threshold for declaring a discovery. It was incredibly strict. To prevent false positives in this ocean of data, an association had to clear a genome wide significance P value of 5 times 10 to the negative eight.
8:04That is a tiny P value. And it required directional consistency. If a transvariant correlated with an increase in a protein in one cohort, but a decrease in another, they just threw the signal out. When you apply filters that rigorous, you expect the surviving data to represent the most fundamental, undeniable architectures of human biology.
8:23And the numbers that emerged completely rewrite our understanding of genetic regulation. Out of all the high confidence protein regulators they mapped, they identified just over 24,000 PQTLs. Here's where it gets really interesting.
8:35The distribution between local and distant regulators. The ratio is just shocking. Only about 5,40 of those associations were local cis variants. And the rest. A staggering 19,698 were distant transpariance.
8:50Wait, really? That's nearly 80%. It is. It proves that the vast majority of protein regulation in the human bloodstream is happening remotely. The systemic network outweighs the local hardware 4 to one.
9:01It forces a total reevaluation of how polygenic our protein levels truly are. The paper illustrates the spectrum beautifully by analyzing the variants explain. Which is just the percentage of a protein's fluctuating levels directly attributable to genetic.
9:15Right. And we see localized extremes, like the protein FCRL 3. Oh, yeah, where genetics explains over 45% of its variants almost entirely driven by heavy local cis regulation. Exactly. But then you have a protein like VEGFR 2, the vascular endothelial growth factor receptor.
9:32Which is critical for forming new blood vessels. Its genetic architecture is sprawling. Its levels rely on a massive interconnected control panel of transregulators scattered all over the genome. And when you map that sprawling network, you don't just find random switches, do you?
9:46No, you uncover the biological pathways the body uses to maintain systemic homeostasis. What's fascinating here is that these distant regulators revealed entirely new biological pathways. Especially regarding how the blood protium is controlled.
10:01Right. The pathway analysis showed transregulators disproportionately mapping back to specific systemic biological processes, most notably N-linked glycosylation. Ah, and linked glycosylation. It's a foundational cellular process.
10:18We're talking about the enzymatic edition of complex glycan molecules, basically sugar chains to the asparagin residues of proteins. It's vital. This modification dictates everything from a protein structural stability and folding to how it interacts with cellular receptors.
10:32And how quickly it is cleared from the bloodstream. Exactly. The data revealed a massive enrichment of transpeak UTLs, mapping directly to the genes encoding the glycosyl transfer ice enzymes. The enzymes that attach those sugars.
10:44Yes. What this tells us is that a mutation in a distant glycosillation gene can subtly alter the sugar tags on 100s of different circulating proteins simultaneously. Oh wow. So a systemic alteration changes how quickly those proteins are degraded or filtered out of the blood.
11:02Right. So the transvariant isn't necessarily changing how much of a target protein is produced. It is altering the global routing protocols that determine how long that protein survives in circulation.
11:12Which is a beautiful realization of systemic biology. But okay, let's transition to the clinical reality here. So what does this all mean for you, the listener? We have a giant spreadsheet of 19,000 distant genetic managers.
11:27How does this actually help patients or change medicine? To understand that, we have to grapple with what the researchers call the discordance problem. And this is a major, somewhat troubling finding for the pharmaceutical industry, isn't it?
11:39It really is. Much of modern drug discovery relies on a statistical framework called Mendelian randomization. Which operates as a kind of natural clinical trial. Since our genetic variants are randomly assigned to conception, they act as instrumental variables.
11:54The theory goes, if you carry a genetic variant that naturally elevates a specific protein, and you also exhibit a higher rate of a certain disease, then the protein must be causally driving the disease.
12:06It aims to strip away confounding variables like diet or lifestyle. However, for Mendelian randomization to be valid, it relies on a strict assumption. The genetic variant must only affect the disease risk through its effect on the biomarker.
12:20But when the researchers compared these genetically predicted disease links against actual observational data from real patients, the convergence was shockingly low. The drop off is severe. They isolated 193 high confidence genetic links, relationships that Mendelian randomization suggested were rock solid.
12:39And how many held up? When cross-referenced with real world biomarker studies, only 52 showed directionally consistent statistically significant support. Barely a quarter. That is a massive mismatch. If we connect this to the bigger picture, we are running directly into the problem of horizontal Piatropy.
12:55Meaning, if you focus solely on a local cis variant to validate a protein as a drug target, you are blinding yourself to the network. Ah, so a genetic variant might increase the expression of protein X, but that same variant might also independently alter a completely different unmeasured pathway that actually causes the disease.
13:16Exactly. Protein X looks like the culprit, but it is merely a bystander reacting to the same underlying genetic stress. So we cannot just rely on the local server in Tokyo to tell us what caused the crash.
13:28We have to map the trans network to verify the true causal pathways. And when we do that, this sprawling map of distant regulators actually start solving major clinical bottlenecks. Yeah, the study outlined specific instances where trans PQTLs provide definitive genetic validation for clinical interventions.
13:47Let's talk about real-world drug repurposing. A prominent example in the study involves rheumatoid arthritis. The researchers identified a very specific circulating protein signature, highly associated with the onset of the disease.
14:01And autoimmune inflammation is notoriously complex, right, often involving dozens of cascading cytokines. It's a mess to untangle. But when they traced the trans PQTLs controlling this specific rheumatoid arthritis signature, the genetic pathways converge decisively on a single distant gene.
14:18TYK too, or tyracine kinees too. Right. TYK 2 is a critical signaling node for immune responses, and pharmaceutical companies have already developed and approved TYK 2 inhibitors, primarily for conditions like psoriasis.
14:32Oh, I see, because this comprehensive genetic map, reveals that the TYK 2 pathway acts as the master distant manager orchestrating the rheumatoid arthritis signature. Yes. It provides robust human genetic evidence for repurposing those existing inhibitors for a new patient population.
14:48It completely bypasses the agonizing trial and error phase of target discovery. And we move from autoimmune inflammation to structural organ degradation with their findings on heart failure. Which is another massive clinical challenge.
15:00Diagnosing heart failure often relies on measuring a circulating biomarker called NT ProBNP. NT ProBNP is released when the muscular walls of the heart are subjected to excessive stretching and stress.
15:14Clinicians use it heavily, but distinguishing whether a biomarker is actively participating in the disease pathway or merely passively leaking from dying tissue is challenging. But the researchers found 7 distinct, highly significant trans PQTLs that physically map the genetic regulation of NT ProBMP directly to establish heart failure risk pathways.
15:36So this systemic mapping upgrades anti-pro BMP. It proves it isn't just a downstream consequence of tissue damage. Exactly. Its distant genetic regulators are inextricably linked to the primary pathology of heart failure.
15:49It definitively supports using it as a predictive biomarker. That's incredible. And the study pushes even further into active therapeutic targeting when analyzing cardiovascular disease and the protein furin.
15:59Furin is fascinating because its traditional role is deeply intracellular. Right. It operates predominantly within the transgolgi network inside the cell. Acting as an enzyme that clees and activates other proteins before they are secreted.
16:13And if you target intracellular furin with a drug, you risk catastrophic side effects because it is essential for fundamental cellular operations. But the researchers triangulated their trans PQTL data with patient survival metrics.
16:27And the genetic architecture painted a vastly different picture of disease causality. Yeah, the convergence showed that the risk of adverse cardiovascular events wasn't being driven by the Furin trapped inside the cell.
16:40The transregulators revealed that it is specifically the extracellular furin. The fraction of the enzyme actively circulating in the bloodstream. That plays the causal role in driving cardiovascular pathology.
16:51This systemic resolution is therapeutic gold. Extracellular furin promotes adverse cardiovascular remodeling. So by mapping the distant regulators, we isolate the specific fraction of the protein causing harm, making circulating fear in a highly viable accessible new candidate for drug targeting.
17:07Completely bypassing the intracellular risks. We're seeing how a systemic perspective unlocks TYK2 inhibitors for rheumatoid arthritis, validates NT probNP for heart failure, and isolates extracellular furin for cardiovascular disease.
17:22It's powerful. But to maintain scientific rigor, we have to acknowledge the boundaries of this specific analysis. Right, the limitations. What did they miss? Well, the limitations are significant and represent the next frontier.
17:35First, the data relies on oling proximity extension assays. While incredibly precise for the 1400 proteins measured, the human body produces tens of thousands of proteins. And countless more distinct isoforms generated by alternative splicing.
17:49We are still mapping a relatively narrow highway system within a massive continental landmass. And the map we do have is heavily skewed. As we mentioned earlier, the nearly 80,000 participants in this meta analysis were overwhelmingly of European ancestry.
18:03Yes. The underlying architecture of these distant transregulatory networks, including linkage to equilibrium structures and allegal frequencies, can differ profoundly across diverse global populations.
18:14So until we achieve this scale of multi-cohort integration across diverse ancestries, Our understanding of the systemic protium remains fundamentally incomplete. Absolutely. Transregulation that dictates disease risk in one population might operate through entirely different genetic nodes in another.
18:34It's a powerful reminder. The precision medicine must be universally inclusive to be genuinely precise. But even with these boundaries. The foundational shift in how we view the genome remains. Yeah. By mapping over 24,000 genetic regulators, this study proves that the vast majority of our blood proteins are controlled by distant genetic managers rather than local ones.
18:56Roughly 80%. It's orchestrated by a sprawling distant network of systemic managers. When we look at these distance signals, we can spot hidden biological pathways like N-linked glycosolation and identify entirely new ways to repurpose existing drugs.
19:10And this raises an important question for you to ponder. What that? What does this mean for the future of clinical trials? If the drugs we are testing actually affect a sprawling, distant network of genes rather than just a single localized target?
19:22Or taking it even further. If environmental factors like a significant shift in your gut microbiome, or exposure to a viral pathogen, can alter the expression of a single distant transcription factor. Could they actively hijack this entire trans network to rewrite our blood prodium in real time?
19:40That is a staggering thought to leave on. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
19:52If you enjoyed this, follow or subscribe in your podcast app, and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
20:06Thanks for listening and join us next time as we explore more science base by base.