A systematic framework to prioritise promoter and UTR variants in 8040 undiagnosed trios from the Genomics England 100,000 Genomes Project, yielding ten likely diagnoses and a validated annotation pipeline for clinical use.
0:00Welcome to Base by Base. Great to be here Today we're looking at some really interesting research. It was published in Genome medicine back in 2025 by Martin Geary and colleagues. That's right. And this paper tackles, well, a persistent challenge in understanding rare diseases.
0:14Which is? It's all about the role genetic variations play, but specifically variations in the non-coding parts of our genes, the promoters, and the untranslated regions or UTRs. Okay, so not the parts that directly code for proteins, which is what most standard tests focus on.
0:30Exactly. And the researchers here focused on people who hadn't actually gotten a genetic diagnosis through those standard tests. So they were still looking for answers. Right, because those promoter in UTR regions, they're like the control switches for the genes, aren't they?
0:46telling them when to turn on, how much to make. Precisely. They're crucial for regulating how genes work. But interpreting variations there. That's been historically tricky. Which means they often get left out of clinical testing.
0:57Yeah. So this study aimed to systematically go after these variations, identify them, analyze them, in a really large group of undiagnosed individuals. How did they approach that? Well, Martin Geary and colleagues developed a specific framework.
1:11The idea was to pinpoint potentially harmful variations, specifically in these regulatory regions, but only for genes already known to cause dominant diseases. Dominant, meaning just one faulty copy is enough to cause the condition.
1:26Correct. And they applied this framework to something called de novo variants. Ah, the new mutations. The ones that pop up in a child, but aren't seen in either parent. Exactly. often strong indicators of a genetic disorder.
1:38They look for these in, um, 840 undiagnosed participants from the Genomics, England 100,000 Genomes project, a big cohort. And it wasn't just about individual mutations, was it? They also did something broader.
1:51That's right. They also conducted what's called a burden test. Okay, what does that involve? It's like looking for an overall pattern. They took aggregated variant data from 7862 unrelated people with rare diseases and compared it to unaffected controls.
2:05The question was basically, do people with rare diseases just have more potentially damaging promoter and UTR variants overall? It's kind of a zoom in, zoom out approach. Yeah, exactly. Looking at individual cases and the bigger picture.
2:19So today we'll break down, you know, how they did it, what they found, and what it all means. Sounds good. Let's start with the methods then. How did they actually go about identifying these genes and regions?
2:30Okay, so 1st things first, they needed a solid list of known disease genes. They turn to panel app for this. Panel app, that's like a curated database linking genes to diseases, right? Exactly. Very reliable.
2:41They picked out 1536 genes known to cause dominant rare diseases with high confidence. And importantly, they made sure these genes had transcripts solicited in the in E1.0 data set. MAA. Yeah, AA provides a sort of standardized high quality set of gene blueprints useful for clinical consistency.
2:59Got it. So they have their disease genes. How do they define those specific non-coding regions, the promoters and UTRs for each one? That sounds precise. It was. For the UTRs, the untranslated regions, they could get the exact coordinates for the Exxons and introns straight from that MAE file, like a map.
3:18Okay. Defining the promoters was a bit more involved. They focused on the proximal promoter, the bit right near where the genes starts being read. The main control switch area. Right. They use data from the encode project, the encyclopedia, DNA elements, massive resource.
3:35They looked for things called candidate cist regulatory elements or CCREs. These are bits of DNA thought to be involved in regulation. Yeah exactly. Encode data tells you about things like chromatin modifications, DNA accessibility, basically, which parts of the DNA are open and likely active in different tissues.
3:51So they're looking for active control regions near the gene start. Pretty much. And because it can vary by tissue. They define both a minimal and a maximal promoter region based on how these CCREs clustered around the transcription starts site, the TSS.
4:04And if there wasn't a CCRE right there. They used a default window, 266 base pairs upstream and 139 downstream of the TSS. Then critically, they use software bed tools to make absolutely sure none of these defined promoter or UTR regions overlapped with the actual protein coding parts.
4:23That's important to keep them separate. Very. This whole process mapped out over 20000000 base pairs of promoter and UTR territory across their target genes. Wow. That is a lot of ground to cover. Now, you mentioned genomics England.
4:36Tell us more about the people involved. Right. The data came from the GEL 100,000 Genomes project. For the De Novo variant hunt. They use GEL version 15 data, looking at 13,913 trios in both parents. And for the burden tests.
4:51That use aggregated data from an earlier version, GELV9, covering over 78,000 participants. Okay, back to those de Novo variants. They start with a big pool. How did they filter down to the most promising ones in these non-coding regions?
5:04They began with high confidence de novo variants already identified by GEL. Then they applied more filters. They focus on those 8040 individuals without a coding diagnosis. They also kicked out. Any variants that showed up, even rarely, in the GEL data itself, were on the big public database, Nome AD, really focusing on the super rare or brand new stuff.
5:24And crucially, crucially, they only kept variants that fell within their defined promoter in UTR regions for those panel app green genes, the high confidence ones, and were relevant to the patient's specific clinical features, their phenotype.
5:38So linking the genetic finding to the actual health problem. Yes. That whole filtering process took them from 1000s down to 1311 potentially relevant to novo variants for just over 1,100 individuals. All right, so they've got the shorter list.
5:52How do they then try to predict which of these might actually be causing trouble? This sounds like a really complex part. It is. This is where annotation comes in, using lots of computational tools. They use ensembles, varying effect predictor, VEP as a queer tool.
6:05VEP tries to guess the impact of a variant. sort of, yeah. And they use plug-ins with VEP too, like Utranitator, specifically for UTRs, obviously, and splice AI, which predicts effects on RNA splicing.
6:17Splicing how the instructions get assembled. Right. And CAD, which gives a general score of how likely a variant is to be damaging. Plus, they looked at conservation scores, Philov, to see if the DNA region is unchanged across species, suggesting importance.
6:31And databases of known variants? Absolutely. They checked against Clinvar, a big database of variance linked to health. Importantly, they excluded anything already known to be benign or likely benign in Clinvar.
6:42So it's like layering different types of evidence using these tools. Exactly. They set specific score thresholds for CAD and file-up to prioritize variants. They use different splice AI thresholds too, depending on whether it was for the de Novo analysis or the burden testing.
6:57And the utranotator. What specific did that look for? It flags things that might mess up translation, like creating a new incorrect start signal and upscream AUG or disrupting existing regulatory sequences in the UTR.
7:10Okay. They also flag variants hitting other known elements like marine binding sites or ires elements, get using CAT and file-up scores to help filter. They looked at RNA binding protein sites, polyand annihilation signals in the 3 foot UTR.
7:22Wow, comprehensive. What about the promoters? For promoters, they looked at known transcription factor binding sites from end code, and they use another tool, Fabian, to predict if a variant within a key part of that binding site, would likely disrupt the factor from binding properly.
7:38That is incredibly detailed. So after all this computational sifting and scoring, what happens? The variants that made it through all that the top candidates then went for clinical review. Meaning a person looked at them.
7:51Yes. Experts compared the patient's actual clinical features using standardized HBO terms, with the known symptoms caused by problems in that specific gene, especially loss of function problems for these dominant genes.
8:04Makes sense. And they needed to check if this whole complex framework was actually reliable, right? Did it work on known variants? Crucial step. They tested it using known pathogenic and known benign variants from Clinvar that fell into their defined regions.
8:17That gives you a measure of, you know, how good is it really at separating the wheat from the chaff. Okay. Now, switching back to that burden test for a moment, how do they set up the comparison groups, cases versus controls?
8:28Right. The cases were genomics, England participants who had at least one relevant green panel app gene linked to their condition, but no coding diagnosis. The controls were unaffected parents from other families in the project.
8:41Why unaffected parents? It's a common control group in rare disease studies. They tried to match them carefully one to one based on sex and genetic ancestry, ensuring no relatives were matched. Do they use all matched pairs?
8:53They refined it a bit. They ended up focusing on 7862 cases who were linked to 100 or fewer relevant dominant genes, and there are 6371 matched controls, mostly of European or South Asian ancestry. This was to try and reduce noise.
9:10And the actual statistical comparison. They used a Fisher's exact test. Basically comparing the frequency of these prioritized variants in cases versus controls, looking across all the different regions and annotation types.
9:22And because they did many tests, they used a Bonferoni correction to adjust the significance level. standard practice for multiple testing. And I saw they also looked at an autism cohort. Yeah, they applied the same framework to De Novo variants from the SFR Simplex collection.
9:38That's an autism research data set. Again, comparing prioritized variants in individuals with autism versus their unaffected siblings. Any other tests? They did do some follow-up lab work on a few GEL individuals.
9:50RNA sequencing to look for splicing issues or changes in gene expression levels, and also DNA methylation tests, which can give clues about gene regulation changes. These were more targeted to dig deeper into specific candidate variants.
10:04Oh, right. That's a really thorough methodology. Let's move on to the results. What did all this work actually uncover? Well, that really strict filtering of the de Novo variants in the 8040 undiagnosed individuals, it initially flagged just 11 candidate variants.
10:17Only 11. That sounds low. It reflects how stringent their filtering was. But here's the interesting part. When clinicians review the patient details for those 11, 9 of them, that's over 80%, were deemed a good match for the person's phenotype, their clinical picture.
10:32Wow, okay, so high relevance among the candidates. Did this lead to actual new diagnoses? Yes, that's the key outcome. Four of those 9 were considered likely new diagnoses. Can you give us the example?
10:44Sure. One was in the 5 foot UTR, the SLC2A1 gene. The variant created a new upstream start code on, potentially disrupting translation, in a patient with symptoms of GLUT1 deficiency syndrome. Okay. Another was a splice disrupting variant near the 5 foot UTR of the NIPBL gene, and someone with features suggesting Cornelia DeLange syndrome.
11:06Then there was a promoter variant, in a conserve spot upstream of ZBTB 18, in a patient with intellectual disability. And finally, a 5 foot UTR splice variant in SCD 5 in a patient initially suspected of having silver Russell syndrome.
11:20And you mentioned methylation earlier. Yes, for that CDD 5 case, follow-up DNA methylation testing actually confirmed a pattern consistent with a city 5 related to order, strengthening the diagnosis. That's fascinating how different lines of evidence converge.
11:34Any other notable finds from that initial group? There was also a potentially interesting cryptic splice variant in the GenAS gene. Or any sequencing showed it led to abnormal splicing, though the exact disease link was still being worked out.
11:46And interestingly, for CTD 5, they later found another person in the wider GEL cohort with a variant at the very same position. This 2nd case also had methylation results supporting the diagnosis. even though the variant didn't quite meet the strictenovo criteria technically.
12:01So adding it all up, what was the final diagnostic yield from these non-coding de Novo variants in this group? In total, including that 2nd Rovity 5 case, they identified likely disease causing promoter or UTR de Novo variants in 10 out of the 8040 individuals.
12:17That works out to about 0.12% of this specific cohort of previously undiagnosed people. Okay, .12%. And what about the reliability check? How did the framework perform on the known Clin bar variants? That showed pretty good results, especially for specificity.
12:30It correctly prioritized about 54%, 66 out of 123 of the known pathogenic variants in their regions. But crucially, it only flagged a tiny fraction less than one%, 24 out of over 3300 of the known benign variants.
12:44So it's good at not raising false alarms, which is really important clinically. Exactly. High specificity. Now, what about the big picture? The burden testing on the larger cohort. Did they find that overall excess of potentially bad non-coding variants in the rare disease group?
13:00Well, that was one of the more surprising findings, perhaps? No, they didn't find a statistically significant enrichment in the cases compared to the controls, not for any specific region type, not for any annotation category, not even when they lumped everything together.
13:15Huh. Despite finding those individual diagnoses, why might that be? The author suggest a couple of main reasons. One is potentially statistical power. Even with 1000s of people. The number of truly causal variants in these regions for any given gene might be quite small.
13:31Detecting a small overall increase might just require even larger coboards. Their power calculations indicated this. The other point they raised was about the control group, the unaffected parents. It's possible this group carries more underlying risk variants than the general population, potentially diluting the signal or making it harder to see a clear difference compared to the cases.
13:52That makes sense. What about the autism cohort analysis? Same result. Pretty much, yes. Applying the framework to the SFRI autism data also didn't show a significant enrichment of these prioritized non-coding variants in the individuals with autism compared to their unaffected siblings.
14:10Okay, so let's pull this together in the discussion. What are the main takeaways Martin Geary and colleagues drew from all this? Well, the big positive is that they clearly show their systematic approach can identify disease causing variants in UTRs and promoters, leading to new diagnoses for families.
14:27That's a definite win. But they also acknowledged that the overall increase in diagnostic yield from these specific regions in this particular cohort seemed relatively modest, that .one, 2% figure. And their choice to focus just on proximal promoters and UTRs.
14:40They argued that was justified, because these regions have a clearer, more direct link to gene function. And we have better tools, relatively speaking, to predict variant effects there compared to, say, very distant enhancers, which can be super complex and tissue specific.
14:57And they highlighted the specificity again. Yes, they stress that the frameworks ability to filter out most benign variants is key for practical application, avoiding information overload. The Klinvar test supported that.
15:10They focused heavily on de Novo variants. What about inherited ones? They pointed out that while de Novo variants were the focus here, due to their high impact likelihood and dominant diseases, the actual annotation and prioritization strategy they built could definitely be applied to inherited variants too.
15:26And the lack of signal in the burden test. What's the thinking there? Their conclusion was that finding a collective signal from individually rare non-coding variants might just need massive sample sizes.
15:37And again, that caveat about the control group potentially carrying a background level of risk variance. Were they upfront about other limitations too? Oh yes, very much so. They noted the limitations like focusing only on genes already in diagnostic panels, mostly using just one main transcript per gene, the main and one, which might miss things.
15:56And the fact that our understanding is incomplete. Absolutely. Our knowledge of non-coding regulation is still growing. And the prediction tools aren't perfect. Their strict filtering might have missed some real ones.
16:08They specifically mentioned lower sensitivity for promoter variants in the Klinvar test as an area needing better tools. And of course, diagnoses might have been made through other means in GEL since their analysis snapshot.
16:19Still, they did have that clear clinical success story. Right. That SLC2A1 variant finding for the patient with GLUT1 deficiency. Since that's a treatable condition. It really highlights the potential clinical value here, and the importance of data sharing to enable reanalysis and find these answers.
16:36So where do we go from here? What did they suggest for future research? Key areas are improving promoter variant prediction tools. Definitely. Also, expanding the analysis to include alternative gene transcripts and incorporate tissue specific regulatory data.
16:52And generally bringing in more well-documented regulatory elements as we learn more about them. Okay, so wrapping it all up. What's the final conclusion from Martin Geary and colleagues? I think the main point is that even with our current, still evolving understanding of the non-coding genome, a systematic analysis of promoters and UTRs using the tools we have can yield important diagnoses.
17:15It's worth doing. And finding more of these non-coding culprits will not only help patients, but also deepen our fundamental understanding of gene regulation. They see their framework as a solid foundation for this kind of work.
17:27Something that can be built upon as our knowledge expands. A valuable contribution then, and for listeners who want to read the original paper, this research was based on an open access article. That's right, published in Genomedicine under the Creative Commons attribution 4.0 international license.
17:41And we'll put the DOI and a link to that license in the description for this program.