LASI-DAD 30× whole-genome sequencing of 2,680 Indian participants produced a 69.5M-variant LD panel that improves genotype imputation accuracy and PRS performance for Indian populations.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So imagine for a 2nd that you are, you know, tasked with building the medical treatments of tomorrow.
0:14That's quite the responsibility. Right, but you're developing these highly targeted precision therapies based on a comprehensive, quote unquote, global map of human genetics. Okay, I see where you're going with this.
0:27Yeah, so you're designing algorithms that will look at a patient's DNA, right? And predict their risk for heart disease or Alzheimer's or diabetes. But there's this, um, this staggering paradox right at the center of your map.
0:40The diagnostic tools you're building completely ignore a quarter of the world's population. Yeah, it really sounds like a premise for a, you know, a dystopian science fiction story. But when you look at the data sets powering modern clinical genetics, well, that missing quarter of the world is a documented reality.
0:56It's wild. And we're looking at a stack of sources today, detailing exactly how big this blind spot really is, and, more importantly, how researchers are finally trying to fix it. Which is a monumental task to say the least.
1:09Absolutely. Because consider this. India is home to over 1400000000 people. We are talking about more than 4500 distinct anthropological groups. Yet historically, the Indian population is made up less than 2% of the participants in global genetic studies.
1:25Less than 2%. Less than 2%. So, the mission for this deep dive is to figure out what really happens when clinical predictions, you know, decisions about your health, your risk for disease are made using data that doesn't represent the patient sitting in the doctor's office.
1:39And the downstream effects of that missing data, I mean, they're profound, to even begin addressing it requires an incredible amount of computational and biological heavy lifting. So today we celebrate the work of the LSI dad team, the University of Michigan, and the Penn Neurodegeneration Genomic Center, who have advanced our understanding of Indian population genetics with their landmark 2026 paper.
2:03Yeah, they took on what is basically the monumental task of building the largest, most nationally representative genetic reference panel for India to date. Which is huge. It really is. But um, to grasp why this requires such a dedicated effort.
2:17We kind of have to look at how the genetic makeup of the Indian subcontinent is actually struck, right? Right, because it's not just one homogeneous group. Exactly. The source material outlines this history of two deeply divergent ancestral groups mixing thousands of years ago.
2:34Yeah. So we are looking at the ancestral North Indians, which is often abbreviated as ANI, and the ancestral South Indians, or ASI. And these groups have very different backgrounds, don't they? Oh, completely.
2:46The ancestral North Indian Group shares ancient genetic roots with populations in Central Asia, the Middle East, and West Eurasia. But in contrast, the ancestral South Indian component is completely unique to the subcontinent.
2:58It's distantly related only to the indigenous populations of the Andaman Islands. Okay, so you have this ancient mixing of 2 distinct groups, which naturally creates a baseline of incredible genetic variation.
3:09But reading through the demographic history, the plot thickens, right? Because of how the population behaves socially over the millennia. Yes, the social structure played a massive role here. Because the sources talk extensively about the founder effect.
3:21And usually, you know, when we hear about the founder effect in biology, it involves a small group of people, physically migrating and getting stranded. Right, like their boat washes up on a remote island, that classic textbook example.
3:33Exactly. But here it's different. It is, because the biological mechanism of a founder event is really all about statistical probability. If you have a massive population, a rare genetic mutation might exist in, say, one in a million people.
3:49Very rare. Right. But if 50 people leave to start a new isolated colony, and just one of them carries that rare mutation. Suddenly, that mutation is present in 2% of the new population. Exactly. And as that isolated group grows over centuries, the rare variant amplifies and becomes highly concentrated.
4:08But the demographic data for India shows that the isolation wasn't driven by physical islands or, you know, impassable mountain ranges. It was driven by social structure. Which is fascinating from a genetic standpoint.
4:19Right. For 1000s of years, there has been this strict practice of indogamous marriages, meaning people marrying almost exclusively within their specific caste, religion, or community group. And the genomic data shows that those social boundaries acted almost exactly like physical ocean barriers.
4:36Wow. Yeah, because of that strict endogamy, the Indian population experienced extreme founder events. But wait, if these various groups have lived side by side in the same cities and villages for centuries, I mean, the DNA hasn't just blended together into one uniform mix over time.
4:52You would think so, but no. The data reveals that the blending was largely halted. The genetic isolation is actually more severe than what we see in Ashkenazi Jewish or Finnish populations. And those are typically the classic textbook models for isolated genetic goboards, right?
5:07Exactly. But here, you essentially have 1000s of distinct, genetically isolated subgroups living in close geographic proximity. Okay, so let me make sure I'm translating the medical implication of this correctly for you, though, listener.
5:19Because these groups have been in these closed genetic loops for millennia, certain rare genetic variants, maybe a mutation that causes a specific neurodegenerative disease, or one that actively protects a person from it, can become highly concentrated in specific communities.
5:35That mechanism is the core of the issue. And it brings us directly to why the current state of global genomic medicine is just failing this population. Because of the GWS databases. Exactly. Modern diagnostics and drug development rely heavily on genome wide association studies, or GWS.
5:54These studies use massive computers to scan the genomes of 100s of 1000s of people, looking for statistical links between a specific genetic variant and a specific disease. The catch being that the database is powering those global studies are overwhelmingly populated by people of European descent.
6:10Overwhelmingly. So if a disease causing variant is highly concentrated in the specific endogamous community in Southeast Asia, but practically non-existent in Europe, our current global diagnostic algorithms will just never find it.
6:24The tool is essentially blind. Illuminating that blind spot is exactly what the LSI dad studies set out to do. The researchers didn't want to just rely on existing fragmented data sets. So what do they do instead?
6:36They pulled high covered whole genome sequencing data from 2680 older adults across 18 states and union territories in India. And when the methodology specifies high coverage, 30 X depth, that implies a level of rigor we should probably unpack.
6:52Because they didn't just skim the genetic code, right? They sequenced every single base pair. Every letter of the genome an average of 30 times over. Yes, 30 times. Why do they need to read the exact same genome 30 times?
7:03That sounds exhausting. Well, it's about eliminating mechanical error. When a sequencing machine reads 1000000000s of base pairs, you know, it makes mistakes. Like a typo. Exactly. If the machine misreads a single letter once, it looks like a brand new rare mutation.
7:16But if the machine reads that same spot 30 times and 15 times it registers a variant, well, the researchers can mathematically prove it is a genuine biological mutation and not a typo. Ah, so it establishes an incredibly robust, trustworthy baseline.
7:33Exactly. And from that massive, verified sequencing effort, they built two primary clinical tools. A genotype imputation panel, and a linkage to sequilibrium panel. Okay, let's start with the genotype imputation panel, because the sources describe this almost like the predictive auto-complete function on a smartphone.
7:51I love that analogy. It's a highly accurate way to visualize the mathematical process. In an everyday clinical setting, running a full 30 X whole genome sequence for every patient is just far too expensive.
8:01Right. nobody's doing that for a routine checkup. No, instead, doctors often use a cheaper micro array test. But that only captures a scattered fraction of your genetic variance. So, imputation software takes those scattered known letters, and uses a reference panel, essentially the genetic dictionary, to statistically predict and fill in the missing blanks.
8:22But, okay, if I start typing a phrase in Hindi into my phone and my phone's autocomplete is programmed with a strictly English dictionary, it's going to fill in the blanks with absolute nonsense. Right.
8:33It's gonna guess completely wrong words. Yeah. The algorithm works fine, but the reference material is wrong. So to accurately auto-complete an Indian genome, the software needs a dictionary built from Indian genetics.
8:46The underlying math demands it. Before the study, the quote unquote dictionaries available for South Asian genomes were either far too small to catch those rare variants, or they were heavily skewed by global data that just didn't reflect the deep localized structure of those endogamous populations we just discussed.
9:02Okay, so that covers the imputation panel. The 2nd tool they built is a linkage to equilibrium, or LD panel. Now, the outline defines LD as a measure of how often certain genetic variants are inherited together, but let's dig into the biology of why that actually happens.
9:20Like, why do certain genetic mutations always seem to travel as a package deal? Well, it comes down to myosis, which is the cellular process that creates egg and sperm cells. Before a cell divides to pass on its DNA, the paired chromosomes physically cross over each other and swap pieces of genetic material.
9:39Right, okay. Think of it like taking 2 separate decks of playing cards, cutting them in half and shuffling the opposing halves together. So the genome is constantly shuffling itself in every generation, which is, you know, what creates genetic diversity in the 1st place.
9:51It is shuffling, but it doesn't shuffle every single individual card perfectly. What do you mean? The cutting process tends to happen at specific physical locations on the chromosome, which we call recombination hotspots.
10:03So pieces of DNA that sit perfectly between those hotspots cards that are physically right next to each other in the deck are highly unlikely to be separated when the deck is cut. They get passed down from generation to generation as an intact block.
10:16Ah, so linkage to equilibrium is simply the statistical measurement of those intact blocks. By mapping them, researchers can say, if a patient has variant A, we are 99% sure they also have variant B, because they are sitting so close together on the chromosome that the cellular shuffling process almost never separates them.
10:36Precisely. And the researchers use 2 specific computational tools to Mac these inherited blocks. L Detect and Big LD. And they serve really complementary functions. How so? Well, LDDec analyzes the genome from a macro level.
10:50It identifies where those major recombination hotspots are the places where the chromosome is most likely to cut the deck, so to speak. Okay, so it sets the outer boundaries. Exactly. Once those bandies are established, Big LD goes inside those large blocks and uses block partitioning algorithms to find the micro patterns, the tiny sub blocks of variants that stick together incredibly tightly.
11:09I see. They then calculated a varal B score, which mathematically quantifies how much the architecture of these genetic packages differs when you compare the Indian population to European or African populations.
11:22So they aren't just looking at isolated single mutations. They are mapping out the entire physical architecture of how DNA is passed down across the subcontinent. That's exactly what they're doing. And the sheer volume of data they uncovered by doing this is staggering.
11:37The new Lasai dad LD panel includes 69.50000000 variants. And to contextualize that number, it represents 170% increase in coverage compared to the previous gold standard reference, which was the 1000 Genomes Project South Asian panel.
11:53Wow, 170% increase. Yeah. The researchers essentially unlocked 1000000s of rare variants that had previously been totally invisible to science. Because the older panel simply didn't sequence enough people or, you know, didn't sequence them at a grip enough coverage to push those mutations above the statistical noise.
12:09Exactly. And when they put this new panel to the test using that auto-complete imputation software we discussed. The performance leap was massive. It really was. Compared to the global top P med panel, which relies heavily on European ancestry data.
12:24This new Indian specific panel, improved infutation accuracy by up to 101%. The mean improvement across the board was 38%. And that jump inaccuracy. I mean, it transforms the viability of clinical diagnostics for this population, but the paper takes it a step further to explore what happens when we try to apply global data the other way around.
12:46They looked at polygenic risk scores or PRF. Okay, let's clearly define PRS for the listener, because this is where all of this background data actually touches the patient in the clinic, right? Very much so.
12:57A polygenic risk score. It takes 1000000s of data points across your genome and calculates a single cumulative score, predicting your genetic risk for conditions like coronary artery disease or predicting complex physical traits like height and body mass index.
13:10Exactly. And the researchers wanted to see what happens when they take a risk score built entirely on European data and use it to predict the height and BMI of the Indian individuals in the LA side dad study.
13:21And the results were a stark demonstration of why localized data is non-negotiable, right? The transferability of those European risk scores just dropped off dramatically, depending on where the patient was located within India.
13:34It did. And I know you're looking at the numbers. And this ties directly back to the ancestral North Indian and ancestral South Indian mixing we mapped out at the start of the deep dive. The genetic history directly dictates the clinical outcome.
13:47Because the European-based risk scores work best for individuals from North India, predicting height with about 17% accuracy. Which aligns perfectly with the biology because the ancestral North Indian component shares ancient genetic ties with Eurasian populations.
14:02So the algorithms recognize the underlying patterns. But even at its peak. I mean, a 17% accuracy rate is extremely low compared to the predictive power those same diagnostic tests have when used on actual European patients.
14:14Oh, it is a fraction of the predictive power. And the accuracy continues to degrade the further south or east you analyze. When they applied the European score to individuals in South India who harbor more of the unique ancestral South Indian component.
14:28The prediction accuracy dropped to 14%. For individuals in East India. The study mentions they have distinct, quote unquote, out of Klein ancestral admixtures. How does the model perform there? Well, out of cline refers to genetic variation that doesn't fit the simple north to south genetic gradient.
14:45East Indian populations often incorporate additional ancient lineages from Southeast Asia. Because the European-based algorithm had absolutely no reference data for those complex ancestral patterns, the prediction accuracy for East Indians plummeted to just 10%.
15:02I just want to pause and look at the real world human impact of those numbers. If a medical network in New Delhi uses a genomic tool and gets a 17% accuracy rate. But a clinic in Kolkata uses the exact same tool and gets a coin flip 10% accuracy rate, that is a systemic health disparity actively baked into the medical infrastructure.
15:22It's deeply problematic. We are building the future of precision medicine on algorithms that actively fail, people, based on their geographic and ancestral background. They do, but the researchers address that disparity head on.
15:36They offer a computational solution called meta imputation. Meta-imputation. Yeah, the goal isn't to throw away the massive global data sense, but to force the algorithms to consult multiple dictionaries simultaneously.
15:47Ah, okay. The notes detail how they layered the new, highly detailed Indian reference panel with massive existing global panels like Topumed and Genome Asia. But how does the software actually know which dictionary to trust?
16:02The logic behind meta imputation relies on probability waiting. A localized panel gives the software incredible accuracy for rare variants specific to that geographic region. But a global panel provides the sheer statistical power of 1000000s of samples for variants that are common across all humans globally.
16:20By feeding both panels into the software, the algorithm calculates the probability of a variant by weighing both the local context and the massive global baseline. Oh, I see. It anchors the global power with regional precision.
16:32Exactly. And the data shows that this meta-imputation approach boosted the accuracy of finding extremely rare variants by an average of 64%. They are literally finding the needles in the genetic haystack simply by combining the maps.
16:46And by making the LSA dad panel completely open access, the research team is allowing scientists anywhere in the world to download this data and start fine mapping causal variants. They can finally begin constructing highly accurate polygenic risk scores tailored specifically for South Asians.
17:03That's incredible. But um, I do have to challenge the long-term scope of this a bit, though. Okay, let's hear it. We establish at the top of the deep diet that India has over 4500 distinct anthropological and endogamous groups.
17:15Yes. The LSI dad study sequenced 2,680 people. It is a rigorous undertaking, obviously, but biologically speaking, can a sample size of roughly 2600 individuals truly capture the genetic complexity of 1400000000 people divided into 1000s of isolated groups.
17:32The mathematical reality is no. And the researchers are entirely transparent about that limitation in their analysis. They explicitly state that this panel is a foundational framework, not a completed map.
17:44Got it. And one major technical limitation they highlight is that the current panel only analyzes biolic variants. Okay, let's unpack biolic for the listener. That refers to specific locations in the genome where only 2 possible letter variations exist across the entire population, right?
18:02Correct. Like, for example, a specific spot on a chromosome where you either have an A or T. Exactly. Focusing exclusively on biolic variance drastically simplifies the computational load, which is what allowed the researchers to actually process the 1000000s of data points required to build this initial reference panel.
18:19Makes sense. But it means the algorithm completely ignores multi-eleic sites. Right, places where 3 or 4 different letter variations might exist at the exact same location across different people. Yeah, and those multi-illolic sites are incredibly complex to map, but they often hold crucial nuanced information about disease risk.
18:37So the current map is high definition, but it's only displaying a specific type of genetic feature. Exactly. Furthermore, when you look at the sheer scale of the region, South Asia represents 25% of the global human population.
18:50A quarter of the world. a quarter of the world. And the genetic diversity driven by those 1000s of endogmous subpopulations means that the linkage disequilibrium patterns, those blocks of inherited traits, can shift drastically, even between distinct communities living in the exact same geographic region.
19:08Right. So if the inherited patterns shift that rapidly from group to group, then 2680 people only gives us a surface level view. It's like, uh, we finally launched a satellite that can take high resolution photos of the continent, but we still don't have the street level view necessary to navigate the intricate details of every single neighborhood.
19:29That's the perfect way to put it. The authors argued that the necessary next step requires highly localized, large scale sampling. The scientific community needs to build cohorts that plunge deep into those individual endogamous groups to capture the variants that are entirely unique to them.
19:43Yeah, that's going to be a huge undertaking. It is, but what the Alasa dad panel accomplishes is proving that the methodology actually works, and it establishes the baseline mathematical architecture required to process that future data.
19:56Wow. Okay, so synthesizing everything we've pulled from these sources. The genetic diversity of the Indian subcontinent holds 1000000s of undiscovered variants that are completely missed by Eurocentric genetic maps.
20:08By building a massive bespoke reference panel, researchers have radically improved our ability to auto-complete and understand South Asian genomes. They are truly paving the way for global precision medicine.
20:21They are. The transition from relying on biased localized data sets to utilizing truly representative global data is the single most critical hurdle, the field of genomic medicine must overcome to ensure modern therapy's work for all populations.
20:36What does this mean for the future of your own healthcare? If your ancestral background isn't represented in the database. That is the big question we're all left wrestling with. This episode was based on an open access article under the CCBY 4.0 license.
20:50You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
21:03Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening and join us next time as we explore more science, base by base.