PaperCast Base by Base discusses a long-read whole-genome sequencing study of 267 individuals from 63 families that increased detection of structural variants and tandem repeats, resolved complex rearrangements, linked repeat expansions to methylation at FMR1, and estimated rare-variant contributions to ASD heritability.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So imagine trying to read a massive 10,000 page manual on how to build a human brain.
0:14Which is, you know, basically what the human genome is. Right, exactly. But imagine that before you could even read it, all the pages were just shredded into tiny strips, mixed up, and then glued back together.
0:24Sounds like a nightmare to read. Oh, totally. You might spot a few, um, single letter typos here and there, but you would completely miss it if an entire chapter was pasted in backward or just duplicated entirely.
0:36Yeah, you'd have no idea about the larger context. Exactly. And for decades. That is essentially what geneticists have been dealing with. It makes you wonder what really happens when our standard genetic tests completely overlook the complex hidden structures that are deeply folded inside our DNA.
0:54We're talking about genomic dark matter here. Yes. Genomic dark matter. Stretches of our genetic code that have remained essentially invisible to standard science. How could shining a light on these hidden genomic structures change our fundamental understanding of neurodevelopmental conditions like autism?
1:13It's a huge question. It is a profound shift in how we look at human biology. So today, we are going on a deep dive into exactly that question, exploring a breakthrough that is finally pulling back the curtain on this hidden genetic architecture.
1:28Yeah, for a long time, we've been trying to understand complex human traits using a map that we, well, we knew it was incomplete. Because our tools just couldn't read the hardest, most tangled parts of the terrain.
1:38Precisely. We just didn't have the tech. Well, today we celebrate the work of Mortazavi and colleagues at UC San Diego, spanning the Departments of Psychiatry and Computer Science, and the Institute for Genomic Medicine, as well as the Ratty Children's Institute for Genomic Medicine.
1:52They've really advanced our understanding of the genetic architecture of autism. They really have. Okay, let's unpack this. We need to set the stage here. Autism spectrum disorder, or ASD, has been this massive clinical and scientific mystery for, you know, decades.
2:08Oh, absolutely. That ones go in for genetic testing. They are hoping for answers. And so often they leave with absolutely nothing. Yeah, and that is the reality for the vast majority of families. When we look at the genetics of ASD.
2:22I mean, we have made some strides over the last decade. Right. We've linked some rare genetic variants to it. Exactly. We've successfully linked rare variants to ASD. These are things like de novo copy number variants where a chunk of DNA is missing or duplicated, and single nucleotide variants.
2:37Which are those single letter typos in the DNA code we mentioned earlier? Right, the typos. But the catch is incredibly frustrating. Those rare variants, they only explain about 4% of the variants in ASD case status.
2:49Wait, 4%? That is a tiny slice of the pie. It's very small If you are listening to this and you've ever had a loved one go through diagnostic testing, You know how maddening a 4% success rate is. It really is.
3:01Now, common single nucleotide polymorphisms, which are the routine genetic variations we all carry and share across the population, those explain about another 11%. Okay, so that brings us to 15% total.
3:14Right. But that still leaves a massive, glaring portion of the genetic architecture completely unexplained. And the problem, it actually lies in our standard technology. Shredded manual. Yes, a shredded manual.
3:27Standard short red, whole genome sequencing is practically blind to larger structural variants. And by larger, we mean any alteration greater than 50 base pairs, right? Yeah, exactly. Anything over 50 base pairs, and it struggles with variable number tandem repeats, too.
3:41Those are the sequences of DNA that repeat over and over again, like a stutter. Yep, like a stutter in the code. So short reads miss a lot of this. Okay, so we know we are missing the big picture with short reads.
3:52Why have we relied on short reads for so long if we knew it was blinding us to these larger structural variants? Well, until very recently, short read sequencing was just the only fast, cost-effective way to read human DNA at scale.
4:05see. It was just an issue of resources. Pretty much. The technology is incredibly accurate at reading individual letters, but it completely struggles with the complex repetitive geography of the genome.
4:16Because the pieces are too small. Right. If you have a sequence that repeats the exact same 100 letters 50 times, a short read sequencer just gets lost. It can't tell if it's reading the 1st copy or the 40th copy.
4:30That makes total sense. So, the researchers in this deep dive completely shifted the paradigm, right? They used long read, whole genome sequencing to crack this mystery. They did, and it changed everything.
4:42Let's talk about how they made that shift. Getting longer reads sounds great in theory. But how do they actually map this out across a population? So they looked at 267 individuals from 63 families affected by ASU.
4:56Okay, that's a decent sized cohort for this kind of intensive sequencing. Yeah, it included 117 offspring. 76 of whom had an autism diagnosis, and 41 neurotypical controls, along with 126 parents. And to get those long reads, they use some pretty advanced tech.
5:13Very advanced. They utilized heavy-hitting sequencing platforms that process massive intact strands of DNA. Like packed bio hi-fi in Oxford Nanopore. Right, exactly. We're talking about platforms that yield average relinks of 1000s and 1000s of base pairs at a time.
5:30Packed bio was giving them around 11,000 base pairs on average, and nanopore was over 5000. Wow. So instead of 150 letter puzzle pieces, they are working with massive chunks, sometimes over 10,000 letters long that actually retain their original context.
5:46That changes the game entirely. And importantly, they didn't just, you know, throw out the old short read data either. Oh, he kept it. Yeah, they merged this brand new long read data with existing short read structural variant data from those exact same subjects.
6:01To create like an ultimate integrated call set. Precisely. But the real breakthrough, the thing that makes this technology so powerful, is that these long read platforms don't just read the genetic sequence.
6:10What else do they read? They detect 5 methyl cytosine directly. Wait, 5 methyl cytosine? That's an epigenetic marker. Yes, it is. You're saying they can see the epigenetics like the chemical tags that turn jeans on or off while they're reading the base code.
6:26Yes, exactly. They can read the raw genetic sequence and the epigenetic DNA methylation simultaneously. That is incredible. To understand why this is revolutionary, you have to look at how we used to do this.
6:38Historically, to see methylation, you had to use separate, messy conversion assays. I remember reading about bisulfite sequencing. It's really harsh on the samples. It is. You had to treat the DNA with harsh chemicals like bisulfite, which physically degrades the DNA sample just to reveal the epigenetic tags.
6:57And now, now they get the sequence and the regulatory tags in a single pass on a single molecule without destroying the sample. I love that. It's like being able to read the text of a book and simultaneously see all the editors highlight marks and sticky notes telling you which paragraphs to ignore, all without tearing the pages.
7:15That's a perfect analogy. But if this technology can read both sequence and methylation at once, I mean, did that actually help untangle the messy genetics of autism, or did it just give us more noise to sift through?
7:28Oh, it illuminated the dark matter in a way we've never seen before. It doesn't just show you the anomaly. it shows you its context. What kind of numbers are we talking about here? Using long read sequencing boosted the detection of gene disrupting structural variants by 33% and tandem repeats by 38%?
7:45Wow, that's a massive jump. It gets better out of over 44,000 structural variants detected, roughly 60%, that's over 16,000 variants, were completely novel. Wait, 16,000 novel variants? Yep. They were missed entirely by previous short red studies on these exact same people.
8:04Furthermore, 98% of annotated tandem repeat regions were successfully genotyped. Over 16,000 novel variants. That is a staggering amount of hidden information. It really is gold mine. When they started sorting through this newly discovered genomic dark matter.
8:20What kind of specific disruptions were they finding in the ASD cases? Well, the rate of de novo structural variants, meaning new mutations that are present in the child, but not inherited from the parents' baseline genome, jumped from 12% in cases to 14%.
8:37Okay, so finding more spontaneous mutations. Mm, right. And a standout discovery was a mosaic duplication in a gene called STK 33. Okay, let me make sure I have this right. That means it's not a mutation in every single cell of the body, right?
8:51It's more like a patchwork where some cells have the mutation and some don't. That is the textbook definition of mosaicism, yes. And because they use long reads, they could actually phase the data. Phasing means they could map exactly which parent that specific mutated chromosome came from.
9:07Exactly. They proved this duplication happened on the maternal haplo type, the mother's chromosome, but it was only present in about 50% of the individual cells. Oh wow. So this tells us the mutation didn't come from the mother's egg.
9:18Right. It happened very early in embryonic development after fertilization as the cells were just starting to divide. So a cell makes a copying error early on, and every cell that comes from that one single cell carries the error, creating that 50-50 patchwork.
9:34You've got it. But what does a duplication in STK 33 actually do to the body? Well, this specific mutation predicted a bizarre 66 amino acid loop alteration right at the protein structure. Wait, if there's a giant 66 amino acid loop just hanging out where it shouldn't be?
9:52Does that break the protein entirely, or does it cause it to do something rogue and dangerous in the brain? It's usually the former. Proteins are like incredibly precise folded machines. If you suddenly insert a massive loop of raw material into the middle of a machine, it destabilizes the entire folding process.
10:09It just gomes up the works. Exactly. The protein likely misfolds, meaning it can't interact with its target molecules properly. And that wasn't the only structural shocker they found. What else was there?
10:20They also found a massive 20 mega-base balanced rearrangement on chromosome 10. 20 megabases. is a gigantic chunk of DNA. It is massive. This chunk broke and reattached. And in doing so, it completely truncated 2 genes.
10:35CCS, ECR2 and SH3PXD2A. Okay, let's focus on that 2nd one. What happens when a gene like SH3 PXD2A gets truncated or, you know, cut short? So it codes for a scaffold protein that's essential for cellular function, and it is heavily intolerant to loss of function mutations.
10:54Because it holds things together. Right. Scaffold proteins hold other cellular components together so they can interact. If you truncate a scaffold protein. It's like pulling the steel beams out of a skyscraper.
11:06Oh, man. So the whole thing just falls apart. The structural integrity of that cellular pathway collapses, which is deeply detrimental during the highly sensitive process of brain development. I was looking through your notes on the different types of structural variants they found.
11:18And there's an entirely new class of variants that sound like something out of an architectural nightmare. Oh, you mean the nested duplication deletion events? Yes, the nested DUPDEL events. They are fascinating.
11:29The researchers found complex rearrangements where a sequence of DNA is duplicated, but then a deletion happens either inside one of those copies were spanning right across them. It is. For example, they highlighted a tandem duplication deletion in the CDC 42 BPA gene, which is highly expressed in the brain.
11:48You know, I was trying to wrap my head around this. And the best analogy I could come up with is looking at a set of architectural blueprints for a house. Okay let's hear it. It's like the builders accidentally duplicated an entire room, say, the kitchen.
12:01They build Kitchen A and Kitchen B. But then right after building the 2nd kitchen, a demolition crew comes in and randomly knocks down half of the original Kitchen A, and the adjoining wall of Kitchen B.
12:13Huh. Yes. That's a great way to picture it. It's an overlapping chaotic mess. And when you map that chaotic mess using long read technology, You literally see a sawtooth pattern in the data, right? Exactly.
12:25The sequence coverage jumps up to 3 copies for the duplication, the extra kitchen, and then abruptly plunges down to one copy for the nested deletion. Creating a jagged sawtooth visual on the graph. Yeah, a real sawtooth signature.
12:38That brings up a huge question, though. How does a cell make such a specific convoluted mistake? Well, when the researchers analyze the breakpoints of these events, the exact molecular sequence where the DNA broke and joined.
12:52They found different DNA repair mechanisms that play within the same event. Multiple repair mechanisms at once. Yes. The duplication might have been caused by one mechanism, and the deletion by another, like microhomology mediated end joining, or MMEJ, and non-homologous end joining, NHEJ.
13:11Okay, let's translate those mechanisms. What are MMEJ and NHTJ actually doing? Think of them as the cells emergency or repair pathways? Imagine a panicked mechanic trying to fix a snapped serpentine belt in an engine while the car is still running?
13:25That sounds dangerous. It is. The mechanic doesn't have time to order the perfect replacement part and just grab whatever duct tape and wire they have and jam the loose ends together. Just to keep it moving.
13:34That is essentially what non-homologous end joining does. It senses a dangerous break in the DNA strand, panics, and forces the brick and ends back together to prevent the cell from dying. But in that rush, it makes mistakes.
13:48Huge mistakes. It often deletes a few lines of genetic code or duplicates them. creating these messy sawtooth mutations. Okay, so with all these complex overlapping errors happening, how do we distinguish a harmless genetic quark from a true driver of autism?
14:04Because clearly our emergency repair pathways make mistakes in all of us, all the time. And this is exactly where that simultaneous methylation data becomes incredibly powerful. Ah, the epigenetics. Right, because they could read the epigenetic tags at the exact same time as the sequence.
14:20They could look at imprinting and gene silencing to see the functional consequence of the mutation. Can you give an example of that? Sure. They found a deletion in the ADNP2 gene. Using the methylation data, they could definitively prove this deletion was on the maternally expressed allele.
14:35Meaning the active copy of the gene, the one the body was relying on to function was the one that got deleted. Precisely. We all inherit 2 copies of most genes, one from each parent. Sometimes the body naturally silences one copy using epigenetics, that's imprinting.
14:51Okay, I'm with you. Because they had the long reads and the methylation data, they could see that the only working copy of ADNP2 was the one destroyed by the deletion. The backup system was already turned off.
15:01Wow, so there was no safety net. Exactly. And the power of this technology gets even more profound when we look at the FMR one gene. The FMR one gene is closely tied to Fragile X syndrome, right? Yes, it is.
15:13We know that massive expansions of a specific DNA repeat, the CGG repeat in the FMR1 gene, cause fragile X syndrome. Which is a major known genetic contributor to autism. Right. If you have an expansion of over 200 of these repeats, the body aggressively methylates the gene, basically turning it off completely.
15:32But what about people with, um, Gray Zone alleles? These are individuals with 35 to 54 repeats, right? Yes, and if you are listening to this and wondering why a gray zone alleal matters, think about the diagnostic limbo these families are in.
15:44A family gets tested and the doctor says, your child has 45 repeats. That's not the 200 required for fragile X, so that's not the cause. Yet the child still has neurodevelopmental challenges. It's incredibly frustrating.
15:56has to be. And this study shines a massive light on that exact limbo. What they found in females with these gray zone allelees is that the specific, slightly elongated allele was hypermethylated, meaning silenced independently of normal X chromosome in activation. Wait, let's slow down on that because the biology of the X chromosome is tricky.
16:19Females have 2 x chromosomes, and normally one is just randomly bundled up and shut off in every cell, right? That's standard X chromosom in activation. Right. So they don't get a double dose of those genes.
16:30Right. But they found that the hypermethylation, the silencing of this gray zone FMR1 gene wasn't just a passive byproduct of that random shutting down of the entire X chromosome. Oh, really? Yeah. The repeat length itself, even though it was only in the gray zone, was driving the body to actively silence that specific gene on that specific chromosome.
16:50They could completely untangle the global chromosome silencing from the local gene silencing because they had the long reads and the methylation data simultaneously. That is mind blowing. It really shows how you can't just look at the raw sequence, like a static string of letters, you know?
17:05You have to see how the cell is actually interacting with it in real time. Exactly. So when they step back and look at all this data together. The rare structural variants, the tandem repeats, the single nucleotide variants.
17:17What is the total impact on autism heritability? When they built a model combining all of these rare factors, they found that they explained 7.4% of the heritability of ASD. Okay, 7.4. Breaking that down, structural variants contribute 5.7%, single nucleotide variants 4.6%, and tandem repeats 3.2% of the variants.
17:40Wait, my math might be rusty, but if I add those up. Um, 5.7 +4.6 +3.2 Well, it's 13.5%. That's way more than 7.4%. It is, yeah. And that is a very common point of confusion when looking at joint statistical models.
17:56Oh, so they aren't just isolated numbers. Right. These aren't isolated, perfectly separate buckets of risk. They are partial contributions within a shared genetic ecosystem. Because some of these variants might overlap or interact with one another in the same individual, you can't just add them up linearly.
18:13The total unique variants explained by combining them all together, caps at 7.4%. Okay, yeah, makes sense. But the crucial takeaway isn't the math quirk. It's that structural variants are actually contributing a massive chunk of that risk.
18:28I mean, more than the single letter typos we've spent the last decade pouring all our resources into. Wow. So let's zoom out and look at the implications here. The ultimate takeaway is that long read whole genome sequencing allows us to resolve these deeply complex structural variations, figure out exactly which parent they came from, and see their regulatory and epigenetic effects.
18:50All in a single test. All in a single test. It's like replacing a grainy black and white photo with a 4K 3D movie. It is a phenomenal upgrade, and it fundamentally reshapes the landscape of clinical practice, but, you know, we do need to critically assess the limitations of the current study.
19:05Right. Right. Nothing is perfect. What are the limitations? Well, they only had 267 genomes. In the world of modern genetics, where humans have so much natural harmless variation, a sample size of 267 is statistically underpowered.
19:20Underpowered to do what exactly? To confidently link specific newly discovered structural variants to autism, you need a massive background population to prove that a variant is actually harmful, and not just a rare but harmless quirk.
19:35So how many do they actually need? They estimate needing about 500 families to really achieve that statistical power. Okay, but I have a sharper pushback here. I was reading the fine print of the methodology, and they relied entirely on peripheral blood samples for this study.
19:51Yes, they do. Wait, if they're looking at blood to understand a neurodevelopmental condition. And you mentioned earlier that for genes like SKK 33, the MRNA expression was totally undetectable in the blood, doesn't that undermine the whole deep dive?
20:06It's fair question. Why use blood to figure out what's broken in the brain? It's a very valid challenge, and it's something we grapple with constantly in clinical genetics. Here's the reality. We use blood because it is a non-invasive, easily accessible proxy for the genome.
20:22the DNA is the same everywhere. Exactly. The DNA sequence itself, the hard coded blueprint is largely identical across almost all the cells in your body, aside from those weird patchwork mosaic errors we talked about earlier.
20:33So a white blood cell has the same basic manual as a neuron in the prefrontal cortex. Okay, the blueprint is the same, but the building being constructed from that blueprint is totally different. Precisely.
20:44And that difference is driven by transcript expression. A MRNA that tells a cell which genes to actually turn on and use differs vastly by tissue type. So blood isn't telling us the whole story. Right.
20:56While blood is a perfect window for reading the raw DNA sequence and finding those complex structural variants, you are absolutely right that it limits our ability to see the downstream functional effects in the brain.
21:08Because if a gene like STK 33 simply isn't turned on in blood cells, we can't study its RNA there. Exactly. And this perfectly points to the need for larger future brain tissue studies, utilizing postmortem brain banks to really see how these complex structural variants alter the active biology of the brain.
21:27Still, it's a stepping stone. But what a massive stone it is. We are finally getting the tools to read the full blueprint without shredding it first. We are. We've proven the long rate technology works at scale.
21:39We've found the hidden variants, and now we know exactly where the genomic dark matter is hiding. So what's the next step? The next decade will be about connecting those newly visible variants to the physical realities of neurodevelopment.
21:49So let's bring it all home. Long read whole genome sequencing is finally illuminating the dark matter of our DNA, proving that hidden structural complexities, tandem repeats, and simultaneous epigenetic changes are critical puzzle pieces in the genetic architecture of autism.
22:06Beautifully summarized. While our current sample sizes are small, this technology gives us an unprecedented multi-layered map of the genome from a single assay. We are no longer shredding the encyclopedia.
22:19We are reading the whole book complete with the editor's notes in the margins. Yeah, that's exactly what's happening What does this mean for the future of personalized medicine and our understanding of human neurodevelopment?
22:29It means we are moving from a binary world of looking for single broken genes into a highly nuanced understanding of how our genome's vast architecture dynamically shapes who we are. This episode was based on an open access article under the CCBY 4.0 license.
22:46You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
22:58Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.