This episode examines a study that used a modular MPRA to test ~11,656 genomic fragments from T2D- and metabolic trait-associated regions in pancreatic beta cells, comparing upstream vs downstream positions and SCP1 vs INS promoters. The work identifies promoter- and position-dependent regulatory activity and implicates HNF1 motifs in INS promoter-specific effects.
0:00Welcome to Base by Base, the papercast that brings genomics to you, wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. You know, we spend a lot of time looking at the genetic blueprint, like it's a, well, like a straightforward machine.
0:13A coding mutation happens, a protein misfolds, and the disease just shows up. It's a very clean narrative. Right. It's a comforting narrative, certainly. You find the structural typo in the Exxon and you've found the root of the disease, but the reality of modern genomics has proven to be, um, far more evasive than that.
0:31Way more evasive. Because when we step into the realm of complex traits, like widespread metabolic conditions, that clean narrative completely falls apart, we find ourselves staring at this genomic landscape that is predominantly, well, dark matter.
0:47Yeah, the regulatory landscape. Right. We know from these massive population level genetic sweeps that around 90% of the genetic signals linked to complex diseases don't actually live in the genes themselves.
0:58They sit out in the non-coding regions of our DNA. Exactly. And these non-coding regions contain the enhancers and silencers that actually modulate transcription. They dictate you know, the spatial and temporal expression of our genes.
1:10And mapping the function of those regions is arguably the central challenge and functional genomics right now. Which is a huge task. Huge. And to query these non-coding sequences at scale, we rely heavily on high throughput platforms.
1:23Most notably, massively parallel reporter assays or MPRAs. Okay, let's unpack this because this is where the central conflict of today's deep dive really begins. We have these massive collections of non-coding variants that we know are statistically linked to diseases.
1:38So we put them into an NPRA to see if they actually alter gene expression. But what if the laboratory testing environment itself is, well, what if it's fundamentally flawed? Like, what really happens when we test a context dependent regulatory sequence using a completely generic cellular backdrop?
1:55Could we be completely missing the crucial triggers for massive widespread diseases simply because our laboratory assays are too generic? That is the exact core vulnerability in how we've approached functional validation for the last decade.
2:08I mean, we've operated under this legacy assumption that the regulatory code is just modular and universally autonomous. Like it just works no matter what. Exactly. But biology actually operates on a highly specific syntax.
2:22A regulatory element doesn't function in a vacuum. If it isn't integrated into the correct physiological circuit, its activity is essentially invisible to our tools. Which brings us to the team that actually proved this context dependency at scale.
2:35Today we celebrate the work of Adelaide Tovar, Yasihiro Kuno, Kirsten Nashino, Maya Bose, Arushi Varshni, Stephen C.J. Parker, and Jacob O'Kitsman from the University of Michigan, who have advanced our understanding of gene regulation and diabetes.
2:49It's an incredibly rigorous experimental design. And for the listener looking to pull the primary source, we are looking at their paper, using a modular, massively parallel reporter assay to discover context dependent regulatory activity in type 2 diabetes linked non-coding regions.
3:04And just for context, that was accepted on April 2, 2026 to appear in Human Genetics and Genomics Advances, which is a CellPress Partner Journal, published on behalf of the American Society of Human Genetics.
3:17Right. So to really appreciate the problem this team solved, we need to look at how NPRAs have historically been constructed. Yeah let's get into the mechanics of that. So when you synthesize 1000s of candidate regulatory fragments to test in parallel, you have to pair them with a promoter to drive the reporter gene, and the standard practice has almost always been to use generic housekeeping promoters.
3:40Like the minimal promoter. Yeah, or the synthetic supercore promoter, which is called SCP1. The reliance on SCP one stems from this old school rule of thumb in molecular biology. The dogma for a long time was that enhancers are highly flexible, position independent, and promoter independent elements.
3:57So the assumption was just that a true enhancer would loop over and start transcription regardless of its distance or orientation. Right, or the specific identity of the core promoter it was paired with.
4:07It was basically thought of as a universal on button. If we connect this to the bigger picture. That assumption has really clouded our functional understanding of type 2 diabetes, hasn't it? Which is a massive global health issue.
4:18Absolutely. We have 100s of associated genetic signals for T2D in non-coding regions. If those regulatory elements need a highly specific cellular environment to function, testing them with generic promoters might mean we are generating massive amounts of false negatives.
4:35We're totally failing to detect the actual causal variants of the disease. Well, it's kind of like evaluating a dramatic actor by making them read lines in a sterile, empty room. That empty room is the generic SEP one promoter.
4:49You might see if the actor can project their voice, but you aren't seeing their actual capability. No, it's a flat performance. Exactly. But put that same actor on a fully dressed movie set with their costars, the native cell environment, and the performance completely changes based on the context.
5:04That is a perfect analogy. You cannot evaluate physiological function in a biological vacuum. So to fix the stage, the researchers at the University of Michigan, completely redesign the architecture of the assay.
5:15They built a modular NPRA, right? Yes. Instead of testing the DNA snippets on that one sterile stage, they tested the exact same pool of sequences across a grid of distinct environments to see how the transcriptional performance shifted.
5:28They started with a massive library of 11,656 DNA fragments. Wow, almost 12,000. Yeah, derived from over 300 regions previously linked to type 2 diabetes and related metabolic traits. And they put these into a pancreatic beta cell line model, specifically, the 83 to 2 routine rat insulinoma cells.
5:49Wait, a rat cell line for human diabetes variant? Yeah, it's a highly robust beta cell model. The core metabolic and secretory pathways are deeply conserved, so beta cells synthesize and secrete insulin in response to glucose, meaning this cell line naturally possesses the transacting factors where T2D variants are natively active.
6:07Okay, that makes sense. So they have the beta cell arena. But they had to change the specific stages inside that arena, right? They tested the fragments by placing them either upstream or downstream of a reporter gene.
6:18Correct. To test the spatial geometry. And then, furthermore, they paired them with either that generic synthetic promoter SCP1 or a physiologically relevant cell specific human insulin promoter called INS.
6:31And then they just measured the RNA to DNA ratio to see which ones successfully drove gene expression. Exactly. The RNA barcode counts relative to the DNA plasmids that entered the cell, show you the differential boost.
6:42Okay, but let me push back on this for a second. Why use the human insulin promoter specifically? Aren't you heavily biasing, the experiment? Like, aren't you basically stacking the deck to guarantee you'll find diabetes related activity in creating false positives?
6:57It's a really valid question, but no, you're not stacking the deck. The mechanics of the assay prevent that. You are providing the necessary physiological syntax, not forcing a result. The INS promoter alone provides the basal machinery.
7:09It sets the baseline. Right. So when we measure the RNA to DNA ratio, We are looking for the differential boost provided by the fragment. If a variant lacks the required regulatory grammar to synergize with that INS promoter, it won't magically fire just because it's in a beta cell.
7:26It just raises the ceiling so you can see what's actually worth. Exactly. The INS promoter contains specific glucose responsive elements and binding sites for pancreas expressed transcription factors, making it the perfect movie set to use your analogy.
7:40So they ran this across the 4 configurations. Upstream generic, downstream generic, upstream specific downstream specific. Yep. and the results were just staggering. The design of the NPRA completely dictated the results.
7:52Out of the 11,656 fragments. How many were significantly active across all 4 configurations? 18. 18. is wild. Just 18 behaved like traditional, universally active enhancers. Most were uniquely active in just one or 2 setups.
8:08It completely kills the universal enhancer idea. It really does. About 700 fragments showed a strict positional bias. Like, they were split pretty evenly between preferring the upstream or the downstream location.
8:20And the promoter bias was even crazier, wasn't it? Oh, yeah. Another roughly 700 showed a promoter bias with a massive 73.4% of those strongly preferring the cell-specific INS promoter. What's fascinating here is what that tells us about the biology.
8:35The fragments that preferred the INS promoter were incredibly rich in binding motifs for HNF1 and MKX 6.1. Right. And for those who might not know, both of those are crucial transcription factors in pancreatic beta cells.
8:50HNF1 is especially important because rare mutations in its coding sequence actually cause MODY. Well, make sure the onset diabetes of the young. Exactly. Modi is driven by high penetrants, rare mutations that break a single note in the system.
9:02Type 2 diabetes is polygenic, driven by 100s of common subtle variants. But here we see this beautiful genetic convergence, the exact same pathway is hit by rare mutations in MODY and common regulatory variants in T2D.
9:16Here's where it gets really interesting. The researchers didn't just stop at the observation. They did a 2nd targeted screen. Yeah, they engineered a specialized MPRA library. They took specific T2D variants like RS 1635852 near the JAZ F1 gene, and RS 11819995 near the ETS1 gene, and they scrambled or deleted the HNF1 motif entirely.
9:43Just took a molecular scalpel to it. And deleting the motif completely killed the regulatory activity, but, and this is the crucial part. It only killed the activity when paired with the INS promoter in the beta cell model.
9:54Oh wow. Yeah, when they tested those same mutated fragments in human skeletal muscle cells, the LHCNM 2 line, it didn't disrupt the activity at all. Because muscle cells operate in a completely different regulatory logic.
10:06They don't even use that H& F1 pathway. So destroying the binding site does nothing to their baseline. Exactly. The regulatory grammar of the pancreas is just a foreign language in the muscle. And there's an expert insight here regarding the position bias, too.
10:18The position bias actually correlated with specific genome annotations. Right, like where they natively sit in our actual DNA. Yes. Upstream bias was linked to active TSS transcription start sites, while downstream bias was linked to typical enhancer annotations.
10:34So even the spatial arrangement in a lab test. has to respect the endogenous geometry of the human genome. If it naturally sits downstream, you can't just force it upstream and expect it to work. Precisely. You have to respect the physical topology.
10:46So when we look at what this means for research, it's huge. The findings prove that the standard generic way of doing NPRAs might be totally blinding researchers to the actual mechanisms of diseases. It's highly probable that our reliance on generic promoters has systematically filtered out a vast suave of context dependent disease drivers.
11:08We've been looking under the straight lamp because the light is better. Which is, I mean, as someone who follows this stuff, it's so frustrating to think about all the previous studies that might have missed key disease drivers just because they use a generic promoter.
11:20I understand the frustration, but we have to view this through the lens of iterative science. The generic promoters were a necessary 1st step. We had to build the high throughput infrastructure first. And now that we have the scale, we can add the physiological precision.
11:34That's a fair point. To truly map out the genetics of type 2 diabetes. We have to match the testing tools to the specific cell type and promoter syntax relevant to the disease. Without this context, we just throw away valuable data.
11:48Exactly. Though the study did note some limitations and next steps. Right, they couldn't fully evaluate variant positioning within the individual fragments themselves, could they? No, due to technical pairing limitations during the cloning process.
12:02Getting the barcodes to align perfectly to test the internal spatial data was too complex for this run, but knowing that gives us the blueprint for the next generation of NPRA libraries. Because the next step is to design new, specialized libraries, to profile the 100s of newly discovered TTD low sci, ensuring that variant position, orientation, and promoter compatibility are built into the test from day one.
12:25Exactly. It's the only way forward. So what does this all mean? If we sum up the central insight here? The regulatory genome is not just a simple set of universal on and off switches. It operates using a highly context dependent syntax.
12:38Right. It's an entire language. And by tailoring our laboratory assays to use physiologically relevant promoters and cell types, We can finally uncover the hidden genetic mechanics of complex diseases that generic tools simply cannot see.
12:53The context is the master key. Absolutely, which brings us to a really provocative prompt for you, the listener, to mull over. We know we've likely missed signals and diabetes. But what does this mean for our understanding of the 1000s of other complex traits from autoimmune disorders to psychiatric conditions that are currently being studied in labs using generic one size fits all tools.
13:15How much of human health is still hiding in the dark matter, waiting for the right stage. It's a question that should keep the whole field awake at night. Definitely. This episode was based on an open access article under the CCBY4.0 license.
13:28You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app, and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
13:40Now, stay with us for an original track created especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.