A commentary calling for generation of tissue-specific molecular data across diverse ancestries to improve fine-mapping, causal inference, and equitable translation of GWAS findings beyond Eurocentric and blood-focused resources.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So I want you to imagine. Uh, you're trying to solve this massive global jigsaw puzzle.
0:16Okay, a puzzle. I with you. Right, but there's a catch. You dump all the pieces on the table and you suddenly realize that like 80% of the pieces you've been given come from just one single neighborhood, the rest of the world basically missing.
0:28Oh wow. Yeah, you're you're not finishing that puzzle. Exactly. Or, you know, imagine a medical student trying to use a universal medical textbook, but 8 out of 10 chapters only describe the biology of one very specific demographic group.
0:42Right. The 2nd a patient from outside that group walks in that textbook is, well, pretty much useless. totally useless. And that's exactly what we're talking about today. We are trying to build this incredible future of precision medicine.
0:54Tailor therapeutics, right? Yeah, yeah, but our genetic blueprints are overwhelmingly sourced from just one corner of humanity. So what really happens when our medical data isn't as diverse as humanity itself.
1:05Well, what happens is we hit a massive bottleneck. I mean, it completely stalls out our ability to translate these genomic discoveries into actual real world clinical care for everyone. And that brings us to the core of today's deep dive.
1:20Today we celebrate the work of Anna Louisa Aruda, Andrew P. Morris, and Eletheria Zagini from Helmholtz Munich, the technical university of Munich, and the University of Manchester, who have advanced our understanding of equity in human genomics.
1:33Yeah, their work is, it's this incredibly detailed open access commentary. It was recently published insul genomics, and it just takes this disparity head on. It really does. And, you know, reading through it, the shift in focus is really striking to me because usually when science has a data shortage, the instinct is just, uh, scale up, right?
1:52Just collect more data. Exactly. Bigger numbers, larger cohorts. But this team's focus on equity. They pivot the whole objective. It's not just about collecting more data, it's about collecting the right data.
2:02So I have to ask, how does focusing on equity actually shift this paradigm away from just hoarding more data? It's a great question. So to understand why the data we have right now is, you know, the wrong data, we kind of have to look at the main engine of modern genetics, which is the Geno Wide Association study.
2:19Right, GWS. Yeah, ATWS. And look, GWS has been an absolutely incredible tool for what, the last 20 years? Oh, totally. A complete game changer. Right, because by scanning the genomes of these massive populations, researchers can pinpoint specific single nucleotide polymorphisms.
2:36The S&Ps. Yeah, exactly SMPs. They find the ones that pop up more frequently and people who have a certain complex disease. But there's a huge butt coming, isn't there? A massive butt. The data sets that power these global scans are just unbelievably skewed.
2:51The authors actually point out that even when you look at massive bio banks in really diverse countries. Like the US or the UK, right? Yeah, exactly. Even there, the participant pools do not reflect the actual demographics of the general population.
3:05It's just overwhelmingly biased toward individuals of European ancestry. Which is, I mean, it's a huge problem. And the authors mention this isn't just an accident, right? It's driven by this historical lack of engagement.
3:17Yeah, it's a very documented, very real long-standing mistrust. you know, stemming from past medical harms and really exclusionary research practices. Right, which means marginalized communities have just been consistently left out of these foundational databases.
3:33Exactly. And the scientific community hasn't historically done a great job of genuinely engaging to fix that. Okay, but let me play devil's advocate here for a second. I want to push back playfully on this.
3:42Aren't human genetics mostly the same across the board? Like SMP is a SNP. A spelling mistake in the DNA is physically the same mistake no matter where your ancestors are from. Right. So why does this regional dice matter so much in the grand scheme of finding, you know, disease cures?
3:59I love that you brought this up because this is where the biology gets really cool. It all comes down to the architecture of human population genetics. Specifically, this thing called linkage disequilibrium or LD.
4:12Okay, LD. break that down for us. So humanity originated in Africa. Right. And because of that incredibly deep evolutionary history. African ancestry populations actually have a significantly higher level of overall genetic diversity.
4:27Oh, interesting. So when humans migrated out of Africa? You went through a genetic bottleneck. They only took a small fraction of that original genetic variation with them as they moved into Europe and Asia.
4:38Wow, okay. And how does that affect the LD? Well, it fundamentally altered the structure of it? Because DNA isn't passed down as single isolated letters. It gets inherited in these big blocks or chunks.
4:49Like getting a whole paragraph instead of just one word. Exactly. Now, over 1000s and 1000s of generations. recombination events. Basically, the mixing of DNA break those blocks apart. So in older populations with higher genetic diversity, like African ancestry populations, there's just been so much more time for recombination to chop those big blocks into tiny, fine green segments.
5:11Ah, I see. So the chunks are much smaller. Much smaller. The correlation between variants is way lower. But in European ancestry populations, because of that bottleneck, those blocks are still really long.
5:22Okay, so if a GWS flags are region associated with, say, diabetes in a European cohort. That flag chunk is going to be massive. It might contain dozens of highly correlated variants. And statistically.
5:37It's a nightmare to figure out which specific variant is actually causing the disease and which ones are just kind of along for the ride. Oh, wow. So they're just guilty by association. Exactly. But when you bring in multi-ancestry data, especially African ancestry data.
5:50You introduce those much smaller LD blocks into your math. Which clears up the noise. It dramatically reduces the statistical noise. You cross-reference the signals, and suddenly you can localize the actual causal variant with so much more resolution.
6:04So diverse data actually makes the science better for everybody. It's not just a box ticking exercise for inclusion. No, not at all. It is a mathematical requirement. If you want to make precision medicine, actually precise.
6:16That is fascinating. Okay, so finding that specific causal variant is the goal. But the paper points out that finding a GOU signal is, like, barely the starting line. Yeah, it's step one of a very long marathon, because the vast majority of these disease signals, they don't even sit inside the coding regions of our genes.
6:37Right. They're in the non-coding regions, the regulatory element. Exactly. The switches that turn genes on and off. So a variant in a non-coding region might be an enhancer, right? It could be influencing the expression of a gene that's located 1000s of base pairs away.
6:51So GWS is like, hey, there's a problem in this general area, but it leaves you completely blind to the actual biological mechanism. Totally blind You have no idea what gene is being altered or how the cell is actually changing its function.
7:04So how do researchers actually map that functional impact? They have to integrate the GWA's findings with molecular data. They use what are called quantitative trait losi or QTLs. QTLs right. And these basically link genetic variants to specific molecular traits, like how much of a certain protein is around or the transcription levels of a specific gene.
7:27And this is where things get really mathematically heavy, right? With colloquialization analysis? Oh, yeah, it gets very sophisticated. Colloquialization is basically the mathematical bridge. It connects the genetic risk you found in GDLG to the actual molecular function you see in the QTLs.
7:43Okay, so let me try an analogy here to see if I'm tracking this. Let's hear it. So if GWS gives us the broad zip code of where the disease variant is, these molecular QTLs act like the blueprint of the house.
7:54I like that. And then colloquialization comes in and mathematically proves that, yes, the broken light switch on this blueprint is the exact direct cause of the power outage in this specific zip code. That is spot on.
8:06Uses Besian statistics to prove that it's the exact same variant causing both things, not just 2 variants sitting next to each other. Okay, and then the paper talks about Mendelian randomization. Where does that fit in?
8:17So Mendelian randomization is what confirms that flipping that light switch actually causes the blackout. It uses the genetic variant as an instrumental variable. Meaning what, exactly? Meaning it treats the genetic variant, like nature's own randomized controlled trial.
8:32Because your genetics are randomly assigned to conception. Right, totally random. And they aren't generally affected by environmental confounders later in life. So this lets scientists infer a true causal relationship between that molecular trait and the disease outcome.
8:51Wow. So that whole framework isolates the exact therapeutic target. And then a pharmaceutical company can come in and design a drug to fix that one specific light switch. You got it. That's the dream of precision.
9:03But, and this is a massive, but the core roadblock that Aruta, Morris, and Zagini are highlighting is that we can only find the light switch if we actually have the blueprint. Exactly. And when you look at figures one and 2 in their paper.
9:14Oh man, those figures are brutal. They really are. They detail the global availability of this molecular data, and the disparity is just glaring. That Eurocentric dominance we saw in GWAs. It's actually magnified in the molecular omex data.
9:30Wait, really? It's worse. It's worse. Whether you're looking at RNA sequencing for transcriptomics or proteomics, mapping protein levels, or metabolomics. European representation is just this massive block across all the charts.
9:43Yes. Well, African, Asian, and Hispanic ancestries are reduced to these tiny marginal slivers. And the authors give a really powerful, real-world example of how this literally halts scientific progress, the Ugandan EGFR study.
9:56Right, the EGFR study. So EGFR is estimated glamariular filtration rate. It's a really critical clinical marker for kidney function. Super important for diagnosing chronic kidney disease. Exactly. So in this study, researchers looked at a cohort of 3288 individuals from a Ugandan population, and they found this completely novel association for kidney function at what's called the GATM Locust.
10:19Okay, so GidoAS did its job perfectly. Perfectly. It found a brand new vital zip code for kidney disease risk, and it was driven by an African-specific variant. But then they hit a wall. A massive brick wall.
10:33They tried to do the functional analysis. They tried to use collocalization to see what molecular traits this variant was messing with, but that specific genetic variant is monomorphic, meaning it's entirely absent in European ancestral populations.
10:48Oh my gosh. And it's super rare in East Asian populations too, right? Right. And because our molecular QTL resources are almost entirely built on European data. The blueprints just don't exist. They don't exist.
11:00The researchers had the mathematical tools to find the mechanism, but the databases were just empty. That is so frustrating. A novel discovery that could have massive implications for treating kidney disease, and it's just stranded.
11:12Completely stranded because of the data disparity. And you know, the complexity of this data shortage actually goes even deeper than that. Wait, how could it be worse? Well, even when we do manage to collect molecular omex data from diverse populations, like through the top P-med program or the UK biobank researchers, hit another critical hurdle.
11:29The tissue problem. The tissue pro. Okay break that down for me. So, think about where most of our diverse molecular data comes from. It's derived from one highly accessible source, blood. Oh, right, because it's easy to draw blood.
11:44Exactly, but if you're trying to map the regulatory mechanisms of, say, a neurodegenerative disease, like Alzheimer's, or a heart condition, right, how useful is molecular data that's drawn exclusively from blood?
11:57Probably not very useful at all, right? I mean, a brain cell is completely different from a blood cell. It provides an incredibly limited view. Gene regulation is highly, highly tissue specific. A neuron in your brain, a cardiomyoside in your heart, a leucocide in your blood, they all contain the exact same genetic code.
12:14Right, the DNA is identical. But they're chrome and accessibility. The way transcription factors bind, and their ultimate gene expression profiles, radically different. It's like the regulatory networks are speaking completely different languages depending on which organ they're in.
12:28That's a great way to put it. So if a variant only disrupts an enhancer that is active in brain tissue, and you're studying gene expression in a blood sample. You're going to get a false negative. Exactly.
12:39You will completely miss the functional consequence of that variant because that enhancer isn't doing anything in the blood anyway. But wait, aren't there massive consortiums dedicated to mapping these tissue specific things?
12:52Like these high-end projects? Yes, in code and the roadmap Epigenomics Project, they exist, but the paper points out a fatal flaw when it comes to global genetics. Which is? Ancestry information for the tissue specific data in those projects is often entirely unavailable.
13:09You're kidding. Nope. We have the tissue data, but we literally don't know who it came from. which makes it impossible to integrate with multi-ancestry GWS. Wow. Okay, what about the genotech tissue expression project?
13:22GTQs. I know that's a big one. So GTex does include metadata, which makes it a cornerstone resource for this stuff. But when you look at the sample sizes, man, it really highlights the scale of the equity problem.
13:34Get me with the numbers. Okay, so in the GTX version 8 release, they did whole genome sequencing across multiple tissues for 838 individuals. 838. Okay. Out of those 838, there are only 103 African-American participants.
13:51And a mered 12 Asian American participants. Well, out of 838. That's barely anything. And it gets worse because the statistical power problem becomes insurmountable when you stratify those numbers by tissue.
14:03Not every participant donated every type of tissue. Right, of course. So for research is looking specifically at, say, liver tissue or lung tissue, the number of minority samples drops into the single digits.
14:14You can't do anything with a sample size of 4 or five. You really can't. You simply cannot generate the statistical confidence required to define a robust molecular QTL. The background biological noise just completely drowns out the signal.
14:28And the paper also argues that even the broad ancestral categories we do use are masking gaps, right? Oh, absolutely. lumping everyone into a monolithic African ancestry category. completely ignores the immense genomic diversity across the entire African con.
14:43Right. It's a huge continent. And similarly, Asian ancestry in these databases is overwhelmingly driven by East Asian samples. So South Asian populations are just heavily underrepresented. We need a much, much more granular definition of sample ancestry.
14:58Okay, so let's summarize where we are. We have a very clear view of the bottleneck. The foundational data is Eurocentric. The limited diverse data we do have is mostly just blood. And the tissue specific resources lack the statistical power to uncover non-European mechanism.
15:14That's the landscape. So how do we fix it? What does the author's roadmap actually look like? They outline a really multifaceted approach. But the main takeaway is that this requires a total overhaul of research priorities.
15:26The 1st pillar, targeted systemic funding. Which means putting the money where the mouth is. Exactly. Financial resources have to be explicitly allocated by major grant agencies to generate multi-ancestry multi-tissue data.
15:41But that requires moving beyond just funding like analytical pipelines on computers. You have to put serious capital into the physical infrastructure of sample collection. Yes. And I've got to raise a logistical and really an ethical pushback here.
15:55Okay, let's hear it. Collecting primary internal tissues. Right. Brain heart liver, collecting those from marginalized or underrepresented communities worldwide. That sounds incredibly difficult. logistically and ethically.
16:09It is incredibly difficult. Because you can't just fly in, take samples and leave. We talked earlier about past medical harms. How do the author suggest we do this without just repeating that history of extraction?
16:19That is the crucial point. The scientific community has to move completely away from that extractive model and move toward genuine partnership. Right. Building actual trust. Exactly. The authors highlight the absolute necessity of global collaboration.
16:34That means active capacity building within those diverse communities. So training local researchers. Trading local researchers, sharing technological resources, and actually establishing infrastructure, that benefits the community providing the data.
16:47It's not just give us your data. It's let's build this science together. Are there any projects doing this right now? Yeah, the human cell atlas is cited as a really strong example. It's this global initiative aiming to map every cell type in the human body, and they have explicitly prioritized, equitable, inclusive sample collection networks all over the world.
17:09That's amazing. And the paper even talks about an ultimate long-term goal, right? Something about longitudinal data. Yes, the monumental goal, the generation of longitudinal omex data across tissues. Meaning tracking people over time.
17:22Exactly. Tracking the molecular profiles of diverse populations over time to understand how environmental interactions and just the aging process affect tissue specific gene expression. Generating dynamic lifelong biological blueprints rather than just these static snapshots.
17:38Exactly. And to pull all of this together, the central insight from Aruta, Morris and Jeannie is clear. Expanding the GWA is Dragnet is totally insufficient on its own. Because finding the variants doesn't mean anything if we don't know what they do.
17:51Right. Identifying disease associated variants in diverse populations will not lead to equitable clinical outcomes unless we also make a massive global investment in generating the tissue specific molecular data.
18:05We need both the zip code and the blueprint. Integrating diverse blueprints with those sophisticated tools we talked about. Localization and multi-omics. That is the only way to ensure the next generation of precision therapeutics works for everyone.
18:19For the whole global population. Not just a tiny fraction of it. Exactly, which leaves you, the listener, with a final broader paradigm to consider. Yeah. We are currently witnessing an absolute explosion in AI driven drug discovery, right?
18:32Where machine learning models predict therapeutic targets by analyzing these vast genomic databases. But if the artificial intelligence systems designing the next generation of therapeutics are training almost exclusively on European centric molecular blueprints.
18:47Are we inadvertently hardwiring a biological blind spot into the very algorithms that are supposed to represent the future of medicine? What does this mean for the future of your own healthcare and how we define what normal human biology actually looks like?
19:00That is a wild thought to leave off on. It really makes you think. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
19:13If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star ratum. If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode, and inspired by the article you've just heard about.
19:28Thanks for listening and join us next time as we explore more science base by base.