This episode examines Deng et al.'s open-access resource that maps 148,198 functionally tested non-coding regulatory elements in human neural stem cells and introduces BRAIN-MAGNET, a CNN that predicts enhancer activity from sequence and pinpoints nucleotides critical for function to prioritize disease-relevant variants.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So let's start this deep dive with a number that's, well, it's pretty startling when you think about it.
0:13Did you know that 98% of your DNA doesn't actually code for proteins? It's incredible for decades. We sort of called it genetic dark matter. Or even junk DNA. But we know now that this massive region, this non-coding part of the genome, it harbors most of the variations that are linked to human disease.
0:31Right. And if you think about, um, neurodevelopmental disorders, so conditions that affect how the brain forms and wires itself up. When clinicians are looking for a cause, they focus on those traditional protein coding genes.
0:46Standard practice. And they only find an answer about 30 to 50% of the time. Which is a huge diagnostic gap. It's huge. It means for more than half of the individuals, the answer is just hidden. It's likely buried in that so-called dark matter.
1:00So the real challenge here isn't just finding these genetic variants. It's about understanding the, I guess, the regulatory grammar of it all. How can we possibly shine a light on these sequences to find the missing answers, especially for something as complex as the brain?
1:15Well, the research we're diving into today offers a really stunning two-part answer. First, they built a massive atlas of experimentally measured function. And second, they use that atlas to train a cutting edge AI to pinpoint the exact sequence change that could be responsible for a disease.
1:33It really feels like a transition point in genomics. It is. We're moving from just mapping the genome to, you know, functionally decoding it letter by letter. Absolutely. And before we get into the nitty-gritty of how they did this, we really want to give a special recognition here.
1:47Yes. Today we celebrate the work of Ruizi Deng, Elena Perentoller, Toss and Stefan Baracott, and all of their colleagues from Erasmus MC and the University of Tubingen. Their work has just profoundly advanced our understanding of functional genomics in neural disease.
2:03It's a fantastic paper. It really is. So to really get why this work was so necessary, we need to detail the core problem, which is functional interpretation. We have these studies, genome wide association studies, or DWAs, and they're great, they're fantastic at finding 1000s of statistical links between S and Ps, single nucleotide polymorphisms and different traits.
2:26Like height or diabetes risk or neurological function, exactly. But the problem is that the vast majority of those S&Ts, they land in these non-coding regions. So you have the link, the association, but you don't have the Y.
2:39You don't have the Y. The functional consequence is totally unclear. It's like GWA shows you a 100 different flags flying near a certain gene, but you have no idea which one is actually controlling that gene's activity.
2:50That's perfect analogy. And we're specifically interested in what are called non-coding regulatory elements or NCREs. Think of them as enhancers. Remote control switches for genes. Exactly. They can boost gene transcription from huge distances away.
3:04And we know that when these switches break, it can cause diseases we call enhanceropathies. But we don't screen for them clinically. We don't. The data is just overwhelming, and we haven't had the right tools to reliably interpret what a variant in one of these regions actually does.
3:20Now, a lot of our listeners will have heard of big projects like encode or the epigenome roadmap. Why weren't those data sets enough to solve this? Ah, that's a great question. Those projects were absolutely foundational.
3:33Critical, but they often identified putative NCREs. Putative meaning supposed. Right. They were based on biochemical markers. So they look for specific histone modifications that decorate the DNA. Like H3K 27 AC for an active enhancer a switch that's on.
3:50Exactly. Or H3K for me one for a switch that's poised or ready. But here's the key distinction. Finding that his stone mark proves the potential for function. It tells you the state of the chromatin. But not the function itself.
4:03Precisely. It doesn't actually measure the ability of that sequence to boost transcription. And for clinical interpretation, and especially for training a good AI model, you need direct, measurable quantitative proof of activity.
4:15That makes total sense. You can't train an AI on a maybe. You can't. You need a verified yes or no. Okay, so let's unpack the core methodology that built that verified atlas. The researchers use a technique with a, well, a ridiculously long name.
4:31Right, it is a mouthful. Chromatin immunoprecipitation, coupled to self-transcribing active regulatory regency quencing. Which we thankfully can just call chip starsec. Yes, thank you. Let's break that down a bit, sure.
4:45You can think of it as a massive functional screening system. It's a type of massively parallel reporter assay. Okay. The real innovation here is that they can take nearly 150,000 unique regulatory elements all at the same time, clone each one into a separate reporter plasma, a little circle of DNA, and then put them all into cells.
5:05So each tiny circle of DNA is carrying one enhancer that they want to test. Exactly. And if that clone enhancer is active, it boosts the transcription of a reporter gene. In this case, something like GFP, which makes the cell light up.
5:17Ah, so you can literally see the activity. You can measure it. By measuring the amount of light, they get a direct, quantifiable readout of just how powerful that enhancer is. They're measuring its native boost strength.
5:29And they did this huge screen in neural stem cells or NSCs. Why was that cell type so important for this? Because NSCs are the source. They're the foundational cells that give rise to the entire central nervous system.
5:42The building blocks of the brain. Right. So they're the most relevant model for understanding how the brain develops. And by extension, what goes wrong in neurodevelopmental disorders or NDDs. By focusing the screen here, they created an atlas specifically tuned to the developing brain's regulatory landscape.
5:59And the result was this incredible data set, a functional atlas of nearly 150,000 NCREs, each one ranked by its measured activity in these NSCs. That's the foundation. And that functionally validated map is what made the next step possible.
6:13This is where it gets, I think, truly transformative. They use this atlas as training data to build an AI tool they named brain magnet. It's a great name. is the brain focused artificial intelligence method to analyze genomes for non-coding regulatory element mutation targets.
6:31The AI, which is a convolutional neural network. It essentially learned the precise regulatory language of the NSC genome. It was trained to look at the raw DNA sequence, just the AST, Cs, and Gs, and predict how active that enhancer will be.
6:46But it's the output that really matters for diagnostics, isn't it? It's not just a general prediction of activity. No, and this is crucial. Brain magnet doesn't just give you a single activity score for the whole element.
6:59It calculates what they call contribution scores or key scores for every single nucleotide within that sequence. So think of it like a long legal contract. Most of the words are just context. The cleave score is like an algorithm that points to the single most important word, the one functional hotspot, that if you change it, the whole agreement breaks.
7:18So it gives you that pinpoint accuracy to look at one specific genetic variant and predict its functional impact. With very high confidence, yes. So let's dive into some of the discoveries they made, just by analyzing this functional atlas.
7:30What do they find out about the genes that are regulated by the most active NCREs? They confirmed a really fundamental principle of developmental biology. Okay. The genes that were regulated by these highly active NCREs.
7:44They showed significantly higher expression levels, and maybe more importantly, they were much more intolerant to loss of function mutations. Could you just quickly define that loss of function intolerance score for us, the PLI score?
7:57Of course. The PLI score is a statistical measure. It tells you how often a gene is seen to be broken or mutated in the general population compared to what you'd expect by chance. So a high score means it's a really important gene.
8:09A very important gene. A high peel ice score means you rarely see that gene broken in healthy people, which implies that breaking even one copy is probably lethal or leads to a very severe developmental problem.
8:20And that correlation confirmed that these highly active NCREs are controlling the most critical genes for neural development. It did. They also use this really clever, comparative approach, looking at the same elements in both neural stem cells and embryonic stem cells.
8:35The comparative chip star sec. Right. And this led them to identify something called primed enhancers. But they run the functional screen in both cell types. They found sequences that were very active in the embryonic stem cells.
8:48They lit up the reporter. Yet, when they looked at the chromatin of those same ESCs, those sequences didn't have the typical active histone mark, the H3K 27 XE. Instead, they had HCK 4 methylation. So wait, the functional test showed the switch was on, but the epigenetic marks showed it was only ready.
9:08Exactly. That epigenetic signature is the very definition of a primed enhancer. It's like the cell puts a bookmark there, epigenetically marking it for later activation. So it's planning ahead. It's planning ahead.
9:19As those cells differentiated into neurons and astrosytes, those primed NCREs gained the H3K 27 ac mark and became fully active. It's beautiful evidence that the developmental timing for building a brain is written into the non-coding genome at a very, very early stage.
9:37That's a huge insight into how the genome preplans its own development. But what about the real test for the AI? Did they prove that brain magnet specific prediction scores actually matter biologically?
9:49They did, with some really robust validation experiments. They took some of these NCREs and made targeted edits. They deleted a tiny 30 base pair segment right on the spot where Brain Magnet gave a high CB score.
10:00And what happened? In 16 out of 17 cases, that small precise deletion just tanks the NCRE's activity. Wow. And what if they deleted a region the AI said was unimportant? No effect. Deleting regions with low CB scores had no functional effect at all.
10:14That is the compelling evidence. It is. It confirms brain magnet can successfully pinpoint the single, functionally critical letters inside these long, complex regulatory sequences. We're moving from just statistical association to a causal mechanism.
10:30So let's talk about the clinical implications. starting with fine mapping common diseases from GWA's data. Right. So with the GY's result, you often get a whole cluster of SMPs that are inherited together.
10:43It's called linkage disequilibrium, and it makes it really hard to tease them apart. But applying brain magnet to large sets of known functional S&Ps from neuropsychiatric disorders, it just showed its superiority.
10:54So? The key here is comparison. Brain magnets keep scores were significantly higher for variants that were already known to affect NCR activity compared to nearby variants that were non-functional. And other tools couldn't do that.
11:07When they benchmarked it against other common scoring methods, tools like CAD, Lindsite, and Former, those other tools often fail to distinguish between the functional and the non-functional S&Ps at the same locus.
11:18So brain magnet cuts through the statistical noise by adding direct functional insight that those other methods lack. Do you have an example? I do. At a schizophrenia associated region on chromosome 6, there were tons of associated SMP standard methods flagged a bunch of them, but Brain Magnet gave the single highest predictive score to a SMP called RS 200483.
11:41And this was the exact variant that other labs had already experimentally validated as being the functional cause. That's incredible. It gives researchers an immediate high confidence target. It does, which brings us to the 2nd, and I think most exciting application, solving rare neurogenetic disorders, finding these enhanced eropathies.
12:01Exactly. The team screen data from the Genomics, England, 100,000 genomes project. They were looking for rarer variants that landed on these high CB score motifs within NCREs that were already linked to known disease genes.
12:14And they found something. They found a fascinating case, involved a rare neurodevelopmental disorder, and the gene RAB 7A. I know RB7A. It linked to Charco Marie 2 cisease, right? Yeah. A peripheral nerve disorder.
12:25That's the one. But in this case, they found a heterozygous rare variant in an NCRE, but it was 45,000 bases upstream of the gene. And it was right in a functional hotspot that brain magnet identified.
12:37Right on the money. And their experiments confirmed it. The variant disrupted a key binding site for the transcription factor YY1, and significantly reduced the NCRE's activity in a dish. Okay, so the variant functionally broke the switch.
12:51Did they show this had a real consequence into a living organism? They did. They use zebrafish, which are a great model for vertebrate development. When they introduced the patient specific NCRE mutant, it resulted in reduced expression of a reporter gene specifically in the central merva system.
13:07Wow. It provided this powerful mechanistic genetic diagnosis based entirely on a non-coding variant that had been completely invisible before, a clear-cut case of an enhanceropathy. This really does sound like a game changer for diagnostics, but no tool is perfect.
13:21What were some of the limitations the researchers highlighted? There are a couple of main hurdles. The 1st one is technical. The NCRE activity is measured episomally. Meaning outside its normal spot in the chromosome in that little reporter, blastmid.
13:35Exactly. And while they showed that a lot of their findings do hold up endogenously, that eposomal context can sometimes miss the complexity of the native chromatin environment, you know, loops, boundaries, other things that can tweak activity in a real cell.
13:50So you're measuring the enhancer's power in a vacuum, which is useful, but the cell might have other volume nogs installed in its natural habitat. That's a great way to put it. And the 2nd major hurdle is the clinical interpretation gap.
14:03Brain magnet is incredibly accurate at predicting the impact on NCRE function. But it doesn't directly predict the clinical phenotype, the actual symptoms a patient will have. And why is that distinction so hard to make?
14:16Well, think about it. Delighting a whole gene might cause one severe well-known syndrome. But just disrupting a single enhancer that controls that gene might only cause a partial effect or an effect in only one tissue.
14:28It could even cause a completely new disorder. So the tool gives us the target. But the clinical follow-up. Connecting that target to the symptoms is still a manual process. It has to be charted one case at a time for now.
14:41That's vital context. So let's bring it all together. What does this mean for us and for the future of precision medicine? I mean, brain magnet is this powerful, functionally validated computational tool.
14:53It's built on a massive foundation of real experimental data from the right cell type. Neural stem cells. It moves us past simple statistical correlation. It lets us prioritize disease relevant non-coding variants with high confidence.
15:08It really is like a magnet for finding those critical regulatory needles in the haystack of our genome. Meatles that have been totally inaccessible until now. This whole deep dive really makes you rethink how we classify genetic disease.
15:21And it leaves me with a pretty big question. If NCRE dysfunction can cause diseases that only partially overlap with the genes known syndrome. How many currently distinct separate syndromes diagnosed decades ago are actually just variations of the same underlying enhanceropathy?
15:37Differentiated only by which enhancer is broken? And when it breaks during development, we might be on the verge of radically reorganizing our medical textbooks. This episode was based on an open access article under the CCBY 4.0 license.
15:51You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
16:04Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science based by base.