Rare copy-number variants were called from 22,319 exomes covering early-onset Alzheimer disease, late-onset disease and unaffected controls, then tested gene by gene for a dosage effect. One locus came back with the cleanest signal in the field: at the central 22q11.21 region, deletions appeared only in early-onset cases, including one that arose de novo, while duplications piled up in controls, with late-onset cases sitting in between. Replication in nearly 400,000 further individuals confirmed it, and overexpressing SCARF2, one of the genes in the narrowed interval, increased amyloid-beta uptake in cells - the direction the protective duplications would predict.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Base Buy Base is now on YouTube too, at Base Buy Base, where every episode gets a video with chapters and the full description come subscribe.
0:16Glad to be here, and I'm super excited about what we're digging into today. Yeah, me too. So today our mission is to explore how extremely rare genetic anomalies, specifically like missing or extra chunks of DNA, might actually hold the key to understanding Alzheimer's disease risk and protection.
0:34And we're using a highly detailed 2026 genomics paper as our grounding source for this deep dive. It's a fantastic paper, honestly. Totally. So imagine reading a critical instruction manual for, like, cellular maintenance, right?
0:47Okay I'm with you. Every word dictates how a cell handles stress clears out debris and basically keeps the whole system from collapsing. And most of the time when we hunt for genetic risk factors, we're looking for a single letter typo in that manual.
1:02Just one tiny mistake. Exactly. But consider what happens when an entire page is just accidentally ripped out of the binding. Oh, wow, yeah. That's a huge loss of information. Or conversely, what if a whole page is printed twice, and it's just jammed right in the middle of a critical chapter.
1:18That would definitely change how you read the manual. Right. So what really happens when having an extra seemingly redundant copy of a genetic page actually shields the brain from disease while missing it triggers early cognitive decline.
1:31How could this change the way we look for protective factors in our own biology? It really forces us to completely reevaluate the whole architecture of genetic risk, you know? Yeah, I'd imagine. Like, we're pivoting from looking exclusively at those tiny sequence errors that break a system to examining these massive structural shifts that might accidentally reinforce it.
1:55Today, we celebrate the work of Olivier Knez, Gail Nicholas, and the massive international team from the European and American Consortia, including University Day Royal Normandy, who have advanced our understanding of copy number variants in Alzheimer's disease.
2:08They really did an incredible job pulling this all together. The scale is just massive. It really is. Yeah. Okay, let's unpack this context a bit because looking for these structural changes, these copy number variants or CNVs isn't entirely new in Alzheimer's research, right?
2:25No, not entirely. But applying it to the broader population like this is. Because we already know that a massive structural duplication can cause the disease in very specific, fully penetrant monogenic cases.
2:37Right. What's fascinating here is how we see that beautifully, albeit tragically, demonstrated with the APP gene on chromosome 21. Right, in Down syndrome. Exactly. Individuals who carry an extra copy of Chromosome 21, develop Alzheimer's pathology with essentially full penetrance by age 70.
2:55Wow, essentially full penetrance. That heavy. Yeah, it's because that extra copy drives the overproduction of amyloid proteins. The mechanism there is clear. Duplicate the gene that makes the toxic protein, and, well, you guarantee the disease.
3:08But the vast majority of Alzheimer's cases globally are non-monogenic. They're very complex. And finding these structural missing or duplicated pages in the complex versions of the disease has been notoriously difficult.
3:23Yeah, the primary hurdle is just the sheer scarcity of these events. I mean, when you scan an average human genome, you generally only find about one rare coding copy number variant per person. Wait, really?
3:36Just one per person on average. Yeah, just one. They are incredibly infrequent. So if you rely on a standard population sample or even, you know, a typical cohort of older Alzheimer's patients. The statistical noise would just completely drown out the signal.
3:50Exactly. To actually see a pattern, you need an overwhelming sample size. And crucially, you need to heavily enrich that sample with younger patients. Oh, because the early onset cases have stronger genetic drivers.
4:01You hit the nail on the head. The genetic drivers and early onset cases tend to be much more prominent because they have to be powerful enough to trigger neurodegeneration decades before the normal aging process would.
4:12That structural scarcity makes the methodology of this deep dive so impressive. It really does. They analyzed ex homes using a software tool called canoes. which relies on reed depth to find these copy number variants. Now, I've heard redef described as weighing a book rather than reading it.
4:29That's a great analogy actually. But I want to make sure we're mapping that analogy correctly onto the bioentramatics here. Well, the book weighing analogy holds up quite well when you scale it. In XM sequencing, we aren't looking at the whole genome, right?
4:41We're only capturing the protein coding regions. Right, just the essential chapters. Yeah. So read depth refers to the literal number of sequencing reads that map to a specific region of the DNA. Okay, so how does that show a deletion or duplication?
4:55Well, if we expect, say, 100 reads mapping to a specific gene based on the baseline data, and we only see 50. That drop in weight suggests a deletion. Exactly. And if we see 150 reads, that excess indicates a duplication.
5:10So instead of trying to read the sequence base by base to find a tiny Tycho, you're running an industrial scale across thousands of genomes, looking for a sudden dip or spike in the sheer volume of data mapping to a neighborhood.
5:23Which is highly efficient for XOM data. But finding the structural variant is really only half the battle. Oh, right. They had to figure out what it was actually doing. Right. So the researchers then implemented what's called an integrated loss of function analysis or um, low analysis.
5:38Integrated low puff. Okay, can you break that down for us? Sure. Historically, genetic studies tend to silo their data. They might look at small truncating variants. You know, the tiny errors that tell a cell to prematurely stop reading a gene.
5:53Or in a completely separate study, They might look at the large physical deletions we just discussed. So instead of just looking at one type of genetic error, they combine different kinds of errors to see if a specific gene was being knocked offline across the board.
6:07Yes. Because from the cell's perspective, the mechanism of failure shouldn't really matter. Yeah, that makes total sense. If a gene is silenced by a premature stop code on, or if it's physically deleted from the chromosome entirely, the cell experiences the exact same deficit.
6:23The protein just isn't getting made. That biological reality is exactly what the integrated loft analysis addresses. By pooling the data from both the small truncating variants and the large deletions.
6:34They measure the aggregate loss of function burden on specific genes. It doesn't matter how the gene was knocked offline. It just counts all the offline genes across the cohort. Exactly. It provides a vastly more accurate picture of a gen's functional necessity.
6:47But to run an analysis like that. Looking for something that only happens once per genome on average. The statistical fortress they had to build is staggering. It's massive. Let's look at the actual scale of the data here.
7:01The discovery cohort alone consisted of 22,319 XMs. That's a huge starting point And within that, they had 4,150 early onset Alzheimer cases, meaning symptom onset at 65 years of age or younger. Right, heavily enriched for the early onset.
7:17Yeah. And they compared those to 8,519 late onset cases and 9,650 unaffected controls. But even with over 22,000 X homes, finding ultra rare variants requires serious validation. Because the signals in a discovery cohort can sometimes just be like statistical illusions.
7:34Exactly. To confirm their findings. They brought in a replication cohort that added another 33,977 affected individuals and 362,322 control individuals. Wait, over 360,000 control? Yeah. We are talking about hundreds of thousands of human genomes acting as a massive biological filter.
7:53That is just wild. And when they ran that integrated loss of function analysis on known Alzheimer's genes through that massive filter, the effect sizes were huge. You really were. Like, for the ABC A1 gene, the loss of function odds ratio was 5.77.
8:07Right. And translating an oz ratio of nearly 6 into a clinical reality means we're looking at a severe biological vulnerability. It's not just subtle thing. No, carrying a loss of function in ABCA one isn't just a slight nudge toward cognitive decline.
8:21It is a heavy thumb on the scale. And we see similar weights with other no risk factors in the paper, too. Yeah, for ABC is 7 deletions, the odds ratio was 2.29. And for CTSB, which the study identifies as a strong candidate gene, right?
8:34Not a firmly established one, the odds ratio was 5.03. Exactly. Five point er 3 for CTSB. Confirming these risk factors in known genes basically proves the tool works. Yeah, it validates the whole method.
8:46But here's where it gets really interesting because the true paradigm shift in this deep dive is what they found on Chromosome 22. Oh, absolutely. There is a specific locus, a neighborhood known as 22 Q11.21.
8:59And the data didn't just show a simple binary risk factor. It revealed a perfect biological gradient based on gene dosage. The 22Q 11.21 locust is structurally fascinating. Due to its genetic architecture, specifically the presence of these low copy repeats.
9:17This region is notoriously prone to unequal crossing over during cellular division. Meaning it's structurally fragile. Exactly. It breaks, deletes, or duplicates much more frequently than other regions of the genome.
9:28Oh, okay. And when the researchers isolated the central region of the slopus, the gradient they discovered was absolute. Deletions at this locus aggressively increase the risk of Alzheimer's. And the data on those deletions is incredibly stark.
9:40Like in the Discovery cohort. These deletions were entirely restricted to early onset cases. Not a single control individual, and that discovery cohort carried a deletion in this region. zero. a huge finding.
9:52And one of those early onset deletions, even a Rose de Novo, meaning it wasn't passed down to the family line. It just appeared spontaneously in the patient, probably due to that structural fragility you just mentioned.
10:04Right. And the absence of deletions in the control group underscores just how damaging it is to be missing this genetic page. But the inverse of that equation is where the therapeutic potential lies. Yes, this is the really exciting part.
10:17Because while a deletion triggers early onset risk, having an extra copy of duplication at this exact locus acts as a profound shield, duplications were heavily enriched in the healthy control group. And the late onset cases fell perfectly into the middle of that gradient.
10:34Right. The frequency deletions and duplications in late onset patients, sat right between the extremes of the early onset cases and the healthy controls. It's a true mirror dosage effect. It's just a phenomenal demonstration of gene dosage.
10:48It really is The odds ratio for the protective duplication is bright three, four, and that's supported by a mega announced p value of 5.52 times 10 to the negative seven. A P value with 7 zeros in front of it.
11:01That means the chance of this being a random statistical artifact is effectively non-existent. Basically zero. Yeah. But we all know population data can't prove mechanism, right? Knowing a neighborhood on chromosome 22 is protective doesn't tell us how it protects the brain.
11:16Especially since that region contains multiple genes. Right. So pinpointing the biological mechanism is where the structural data merges with cellular biology. Exactly. And through their integrated analysis, the research team zeroed in on one specific gene within that locus called scar F2.
11:32Scarf2. Yeah, it codes for a stavenger receptor. Basically, it's a protein that sits on the surface of a cell, identifies extracellular debris or toxic proteins, and initiates the process to engulf and clear them.
11:45A microscopic garbage collector. Essentially, yes. But to prove scarf 2 was actually responsible for the protection, they had to take this out of the population data and put it into a Petri dish. Right.
11:54They had to test it in vitro. So they cultured human microglial cells, which act as the brain's primary immune responders, and artificially over expressed the scar F2 gene. Mimicking the state of someone walking around with the protective genetic duplication.
12:08Exactly. Then they flooded the environment with fluorescently tagged amyloid beta, which is that toxic protein known to aggregate an Alzheimer's disease. And the resulting cellular activity was striking.
12:20What happened? Well, the microbial cells with the overexpressed scar F2 demonstrated a massive increase in amyloid beta uptake. The P value there was 2.2 times 10 to the -5. Wow. So the extra copies of the scar 2 instructions directly translated into the cells clearing the toxic protein at a vastly accelerated rate.
12:40Yes, but, you know, there's always a caveat to test for. Yeah, I have to admit, when I 1st read that part, my immediate thought was about general cellular metabolism. Right. Like if you artificially ramp up a scavenger receptor, couldn't you just be inducing a state of hyperphagocytosis?
12:55A feeding frenzy. Exactly. The cells might just be on a microscopic feeding frenzy, blindly eating everything in their path rather than specifically targeting the Alzheimer's pathology. And rigorous biological assays anticipate exactly that kind of mechanical ambiguity.
13:11So how did they prove it wasn't just a feeding frenzy? To prove the specificity of SCARF2, the researchers ran a parallel control experiment. They introduced inert, latex beads, harmless, microscopic pieces of plastic to those same SCARF2 overexpressing cells.
13:28Oh, okay. If the cells were simply in a generalized feeding frenzy, their consumption of the latex beads should have spiked proportionally. But the latex bead uptake did not change at all. The P value for the plastic beads was .11, making it statistically indistinguishable from baseline.
13:44There was no microscopic plastic eating bender. Zero bender. And that unchanged control is the crucial linchpin of the entire experiment, right? It proves that the scarf 2 protein isn't just a generalized vacuum cleaner.
13:56It specifically recognizes and clears the toxic amyloid beta. Exactly. The mechanistic through line is complete. At a population level, a structural duplication at 22Q11.21 provides an extra copy of the SCARF2 gene.
14:10And at a cellular level, more ScarF2 receptors mean a higher capacity to clear amyloid beta. Yep That accelerated clearance prevents the toxic accumulation that drives neurodegeneration, which shields the brain and reflects as that .34 odds ratio in the population data.
14:26The mechanics make total sense, but I feel like we need to be incredibly careful with the word shield. Yes, very careful. Because if someone has an os ratio of .34, their risk is slashed, but it isn't zero.
14:39If I carry this duplication, it does not mean I am completely immune to Alzheimer's, does it? No, not at all. And the authors are highly explicit on this point. The protection is relative, and it operates within a framework of complex determinism.
14:52So it's not an absolute cure. Not an absolute cure. A genetic duplication might build a thicker wall, but if the biological siege is relentless enough, that wall can still be breached. And we see that play out in the data, because some individuals carrying the protected duplication still developed late onset Alzheimer's disease.
15:09Right. How does that happen if their microglial cells are so highly efficient at clearing the amyloid? It really comes down to a biological tug of war. Those individuals who developed the disease despite carrying the SCARF2 duplication, frequently carried other heavy genetic risk factors, such as the APOE4 allele.
15:29Oh, APOE. Yeah. If APOE4 is driving an aggressive accumulation of toxic proteins, the extra scar F2 receptors are working overtime to clear it. But they can only do so much. Exactly. For a long time, the scarf, if 2 duplication might successfully delay the onset of symptoms, pushing back the clock of cognitive decline.
15:49But eventually the sheer volume of the pathology driven by APOE4 can just overwhelm that protective clearance capacity. So the extra garbage trucks are running 2047. But if the factory is producing too much toxic waste, the system still eventually fails.
16:04That's a great way to put it. This complex interplay brings up a vital clinical reality, though. When people hear about a genetic variant that slashes Alzheimer's risk by 2 thirds, the immediate instinct is to ask their doctor for a test.
16:17Which is an entirely understandable impulse, honestly, but one we must firmly caution against. These 22Q, 11.21 deletions and duplications are extremely rare. The vast majority of the global population does not carry either of them.
16:31So people shouldn't be running out to get tested. Exactly. No one should be seeking clinical testing or looking for personal reassurance based on this specific locust right now. The real value of this discovery lies in mapping the mechanics of the disease, and providing a blueprint for drug development, rather than offering a new screening tool for the general public.
16:52That makes sense. But there is one specific subset of the population where this data actually has immediate actionable clinical implications, doesn't it? Yes, and that relates to De George syndrome. Can you explain that connection?
17:04Sure. DiGeorge Syndrome is a condition caused by large chromosomal deletions in this exact 22Q 11 region. Historically, the medical community focused primarily on the immediate severe complications of the syndrome, things like heart defects and severe immune deficiencies, which unfortunately resulted in limited life expectancies.
17:25But advances in pediatric and cardiovascular medicine have completely changed that trajectory, right? Completely. Today, patients with D. George Syndrome are regularly aging into their 50s, 60s, and beyond.
17:37And as that population ages, this study highlights a new critical vulnerability for them. Yes. Because their underlying condition is driven by deletions that often encompass the scar F2, KLHL 22 MED 15 region, this genomic data strongly suggests they're at a significantly elevated risk for early onset Alzheimer's disease.
17:58Wow. So armed with this knowledge, medical professionals should be offering these aging patients proactive, cognitive follow-ups, right? Anticipating the risk rather than just waiting for severe symptoms to manifest.
18:08Exactly. It's a profound translation from a massive data set to individualized patient care. It really is. But, you know, science requires us to look at the boundaries of the data, too. Despite the monumental scale of the replication cohorts, What are the inherent limitations of a study focused on structural variants that are this rare?
18:26Well, we have to acknowledge 3 primary limitations here? First, despite screening hundreds of thousands of exomes, the absolute count of individuals carrying these specific deletions and duplications still remains quite small.
18:39Right, because they're so ultra rare. Exactly. And anytime you're dealing with ultra rare variants, the small sample counts introduce a degree of statistical fragility. Yet when the N is small, the data is susceptible to outsized impacts from minor variables.
18:54Precisely. Second, the intrinsic scarcity of these variants, averaging only one rare coding CNV per genome severely limits the statistical power needed to discover other structural variants across the genome.
19:07Oh, I see. might be other protective ones out there. Right. There may be dozens of other duplicated pages providing subtle protection, but we currently lack the sheer volume of data required to separate them from the background noise.
19:18That makes sense. What's the third limitation? Finally, the genomic analysis in the study was restricted exclusively to individuals of European ancestry. Ah, which is a systemic bottleneck in global genoma.
19:30Your huge one, yeah. Because we know that the genetic background like, the rest of the cellular instruction manual deeply influences how any single variant behaves. Totally So before we can universally apply these findings or design therapeutics around them, this exact structural analysis must be replicated in diverse global populations.
19:49If we connect this to the bigger picture, the intersection of these rare variants with diverse genetic backgrounds will likely reveal an even more intricate web of neurodegenerative risk and resilience.
20:00It reinforces that Alzheimer's isn't simply a single broken gear. It is the sum total of how all the typos, the missing pages, and the extra pages interact over a human lifetime. So what does this all mean?
20:13Well, to summarize, this deep dive reveals that rare structural changes in our DNA, like deletions and duplications play a direct role in non-monogenic Alzheimer's disease, specifically, while a deletion at the 22Q 11 quote 11.21 locus acts as a harsh risk factor, a duplication there acts as a profound shield, potentially by boosting the brain's ability to clear toxic amylode proteins via the SCARF2 gene.
20:38What does this mean for the future of Alzheimer's therapies if we can artificially mimic the protective effects of a genetic duplication? This episode was based on an open access article under the CCBY 4.0 license.
20:49You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
21:00Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science based by base.