Herger et al. present a pooled prime editing platform in haploid human cells that installs and assays thousands of short variants in their endogenous context. Using surrogate targets, co-selection and stringent pegRNA filtering, negative and positive selection screens identify loss-of-function variants in SMARCB1 and MLH1, including non-coding ClinVar variants that alter splicing.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine a massive 3000000000 letter dictionary of human DNA.
0:12Inside, there are 1000000s of slight misspellings. We know some of these typos cause devastating diseases. But the vast majority are total mysteries. What if we had a molecular search and replace tool that could simultaneously test 1000s of these typos in living cells to see exactly what breaks?
0:30How could this change the way we diagnose genetic diseases? And what really happens when we start editing the dark matter of our genome? Well, I mean, it completely redefines our approach to all those genomic mysteries.
0:42You know, instead of just reading that massive DNA dictionary and sort of guessing what the typos might mean, we can now actively rewrite them. And we can do it at an unprecedented scale, observing the biological fallout in real time.
0:53Today we celebrate the work of the research team at the Francis Crick Institute, who advanced our understanding of functional genomics and variant interpretation. Yeah, and their work addresses, um, one of the most pressing, frustrating bottlenecks in modern medicine.
1:07The core scientific problem here really revolves around something called variants of uncertain significance, or VUS's. Oh, VUS's. I feel like anyone who has had a clinical genetic test or, you know, knows someone who has might have actually encountered this term.
1:23Oh, absolutely. You get your results hoping for a definitive answer, and instead, you get a medical shoulder shrug. That is the perfect way to describe it, honestly, because when a patient gets their genome sequenced, doctors check the results against massive global clinical databases, like, uh, Clinvar, for instance.
1:42The right, the big ledger. Exactly. Clinvar acts as a ledger of human genetic variation. But it is just absolutely overflowing with these unclassified variants. I mean, they are genetic mutations that have been discovered in actual patients, but clinicians have no idea if they're just harmless quirks.
1:57Like having curly hair instead of straight. Exactly. Or if they're the underlying cause of a severe life-threatening disease. So doctors are just staring at these typos and don't know the definition. Right.
2:08And historically, scientists have tried to clear up this uh, diagnostic muddy water. using multiplexed assays of variant effect, basically maybes. The idea is to synthesize many variants at once in a lab to see what they do.
2:22Okay, but wait, haven't we been using CRISPR to cut and paste DNA for years. Why is it so hard to test these variants? I feel like I read about a new CRISPR breakthrough every single week. Yeah, that's super common.
2:34The public perception of CRISPR is often that it's this magic wand that can seamlessly fix anything. But traditional crisper, you know, the original cast 9 system is more like a pair of blunt molecular scissors.
2:45Ah, blunt scissors. Yeah, it physically severs both strands of the DNA double helix. And when that happens, the cell panics. It views a double strand break as catastrophic damage. Like the instruction manual just got ripped in half.
2:58Exactly. So it frantically tries to repair it, often making messy, unpredictable mistakes in the process by just like jamming the broken ends back together. Right, so it's excellent at breaking a gene to turn it off entirely, but not necessarily rewriting it perfectly.
3:12Spot on. And even if you try to use a process called homology directed repair or HDR, where you give it a template to copy from? Yeah, where you provide the cell with a physical DNA template to fix the brake.
3:24But the thing is, it's incredibly inefficient. The cell's natural, messy repair pathway usually outpaces the clean template repair. Oh, wow. Plus, this clean repair process is confined to tiny little windows of the genome, maybe like 100 to 200 base pairs at a time.
3:39So if we want to test all the 1000s of mysterious variants scattered across a patient's DNA, we needed a better way to install any short variant anywhere we want without triggering that cellular panic.
3:53Okay, and this is where the research team's approach gets really fascinating. Because they didn't just use standard CRISPR molecular scissors. They use an innovation called a pooled prime editing platform.
4:03Yes. I've read a bit about prime editors. They don't make that blunt double strand cut, do they? That is the key physical difference right there. Prime editing uses a modified cast 9 called a Nick Ace.
4:16Instead of severing the DNA entirely, it acts more like a scalpel, carefully nicking just one single strand of the double helix. Just a tiny nick. Yeah. But the real ingenuity is what that nick ace is attached to, a reverse transcript taste enzyme.
4:31A verse transcript is. That's an enzyme that writes DNA from an RNA template, right? Some viruses actually use that mechanism to copy themselves into our genome. They do, yeah. But here, it's been engineered for precision editing, and it's guided to the exact right spot by a highly specialized RNA molecule called a peg RNA, a prime editing guide, RNA.
4:50Okay, a pig, RNA? The pig RNA is brilliant because it doesn't just find the genomic address like standard CRISPR guides. It literally carries the replacement text on its tail. Oh, wow. So it's a true molecular search and replace function.
5:03It finds the spot, the Nickase nix the single strand open a tiny flab, and the reverse transcript case reads the replacement text on the PegR name, writes it directly into the genome sealing it in. Exactly.
5:13Without ever triggering the panic of a double strand brake. That's incredible. But doing this to 1000s of cells at once, like testing 1000s of different typos in a single pool requires a very specific biological canvas, they used HAP1 cells.
5:27Yes, HAP1 cells are crucial here. Well, let me see if I can put this into an analogy to explain to you listening why this matters. Normally, our cells have 2 copies of every gene, a built-in biochemical backup system, but HAP1 cells are haploid.
5:41They only have one copy. It's genetic high wire walking without a safety net. If we edit that one gene and break it, we see the disaster immediately. I love that analogy. That unmasked visibility is exactly why haploid cells are the perfect model for this kind of deep dive screen.
5:56We get to see the true isolated effect of a single mutation because there's no healthy backup copy to mask the damage. Right. So for this study, the team used HAP1 cells expressing a highly optimized prime editor, and to push the editing efficiency even higher, they temporarily suppressed a protein called MLH1.
6:15Wait, I recognize MLH1. Isn't that the gene responsible for a mismatch repair? Like, why would they want to suppress the cell's own spell checker? Because prime editing fundamentally works by creating a deliberate mismatch for a brief moment.
6:28When the reverse transcript case writes the new DNA sequence, it temporarily doesn't match the unedited strand across from it. I see. So if the cell spell checker is fully active, it will spot that mismatch, assume it's an error, and try to revert the edit back to the original sequence.
6:47Oh, so by suppressing MLH one, they essentially put the spell checker to sleep so then the new prime edits would permanently stick. Exactly. You have to turn off the alarm system while you change the locks.
6:57That makes perfect biological sense. But managing this on a massive scale, I mean, that requires some intense quality control. You can't just dump a massive library of 1000s of peg RNAs onto a plate of cells and hope for the best.
7:10No definitely not. First, they used a technique called co-selection. Yeah, this is where they co-edited a completely different essential gene called ATP1 A1. At the exact same time they made their variant edits.
7:21And that specific ATP one A1 edit makes the cell resistant to a toxin called Uabane. Let's unpack the mechanism there. Why does Oobane kill a normal cell? So Ubane is a highly toxic compound that binds to and jams the sodium potassium pumps on the cell membrane.
7:38These pumps are constantly working to push sodium out and bring potassium in, keeping the cells fluid balance stable. Like bailing water out of a boat. Right. And if you jam those pumps, the cell can't regulate its internal pressure, it swells up with water and literally bursts.
7:53But the prime edit on the ATP1 A1 gene changes the physical shape of that pump just enough that the Uabane toxin can no longer bind to it. So by bathing the entire pool of cells in Uabane, it acts as a biological bancer.
8:06Exactly. Only the cells that successfully took up the prime editing machinery get the pump upgrade so they survive. The rest burst and die. It's an incredibly harsh but effective way to clear out the noise.
8:17It creates a perfectly clean slate. There's another layer of complexity here. Even if the cell took up the machinery and survived the toxin. How do you know the specific Pej RNA for the mysterious variant you're actually trying to study worked.
8:33Because some pejarnays simply fail to make the edit once they get inside the cell. Right, it might just be a dud. This is where the surrogate targets come in. And I have to admit it took me a 2nd to wrap my head around it.
8:43They included a tiny 55 base pair of genetic sequence on the lentivirus itself. And the lentivirus is basically a hollowed out microscopic delivery truck they use to get the genetic material into the cell.
8:55Yeah, exactly. Why edit the vehicle? It's like sending a cheap GPS tracker through a delivery route before sending a priceless artifact. If the tracker never arrives, you know that route is a dud. By reading the surrogate target, we know if the specific pej RNA is actually working before trusting its effect on the real genome.
9:13That distinction is the bedrock of their entire data set. Because think about it. If you skip that surrogate step and observe no change in a cell's behavior, You might incorrectly assume the genetic variant is completely harmless.
9:26Right, a false negative. Exactly. In reality, the edit might never have happened at all. So the surrogate target prevents us from confusing a technical failure with a biological result. By deep sequencing those surrogate targets on the delivery vehicles, they could confidently filter their beta sets.
9:44And only look at the variants that were successfully installed. Okay, so we have the machinery, the biological bouncer, and the GPS tracker. What happens when they actually turn this loose on the genome?
9:53Let's trace their negative selection screen. They targeted a gene called SRCB1. Yeah, so SRCB one is a non-negotiable essential gene for cell survival. It encodes a core piece of a protein complex that remodels chromatin.
10:08Which basically means it helps unwind the tightly spooled DNA so other machinery can read the instructions, right? Spot on, because it's essential. If you break it in these haploid cells, The cell loses its ability to read its own DNA and it dies.
10:21That's negative selection. You apply the edits and you sequence the population over time to see what disappears. They tested over 7500 peg RNAs on this one gene. And because they're using those haploid cells without a safety net, if a peg RNA successfully installs a destructive typo, that cell is doomed. Exactly.
10:42By using their surrogate target filter to 0 in on only the highly efficient peg RNAs, they track the cells over 34 days in culture. The cells carrying the destructive loss of function edits were severely depleted, and they confidently identified 12 of these significantly depleted variants.
10:58One of those variants really caught my eye because of where it was located. It was an intronic variant. It wasn't even in the main coding part of the gene. It was buried deep in the intron. You know, the non-coding intervening space between the genetic instructions.
11:10It was listed in the clinical databases as a total mystery. Yeah, and that specific prime editing screen proved it caused catastrophic splicing errors. Let's zoom in on that splicing process for a second.
11:21When a cell reads a gene, it first creates a rough RNA copy that includes both the vital instructions and the Entronic filler. Think of introns like commercial breaks in a movie broadcast. Before the cell can use the RNA to beta protein, a machine called the Splice of Some has to perfectly cut out those commercial breaks and paste the movie scenes back together so it plays smoothly.
11:44Oh, like that. This intronic variant acted like a corrupted signal. It destroyed the cellular queue that tells the splices some where to make the cut. So the spicy sum gets confused and leaves a chunk of the commercial break right in the middle of a crucial movie scene.
11:58The resulting protein is a mangled, useless mess, and the cell dies. It just beautifully illustrates why we cannot afford to ignore the regions outside the main coding sequence. Which transitions us perfectly to their positive selection screen, because finding the variants that kill a cell is one thing, but how do you track down the variants that allow a cell to survive when it shouldn't?
12:21For this, they targeted the MLH1 gene, which is that mismatch repair spell checker we discussed earlier. Right. Breaking MLH1 in humans is really dangerous. It increases cancer risk because genetic errors start piling up.
12:32But in a Petri dish, breaking it provides this strange advantage against a specific toxic drug called 6th thioguine, or 6TG. Yeah, the mechanism here relies on the cell's own quality control. 6TG basically mimics a normal building block of DNA.
12:49So when a healthy cell incorporates it. The MLH1 mismatch repair system spots it as a fraudulent molecule. Like a counterfeit bill. Exactly. The system tries to cut it out but fails, eventually leaving a toxic gap in the DNA.
13:01The cell realizes its genome is irreparably compromised and triggers apoptosis, which is programmed cell death. But if we use prime editing to break the MLH1 gene first. The cell is functionally blind to the drug.
13:13The spell checker is broken, so it never spots the fraudulent 6TG, never tries to fix it and never triggers cell death. It just ignores the drug entirely and keeps multiplying. It is the ultimate survival of the unfittest.
13:24When they apply the 60G drug to the pool, every healthy cell triggers his own death. The only cells that grow out are the ones where the prime editor successfully installed a destructive typo in the MLH1 gene.
13:37Wow, and they use this positive selection pressure to screen a massive 60 kilo-based non-coding region around the gene, testing nearly 900 variants from the Clinvar database. Yeah, massive scale. These are the mysterious typos scattered in the vast stretches of DNA surrounding the gene.
13:55The scale of this is wild. Imagine a 60,000 letter non-coding sequence. Finding a single broken letter in that sea of genetic text that actually causes disease. It completely shatters the outdated concept of junk DNA.
14:09This non-coding dark matter is actively regulating our health, and a typo deep in the margins can just shut down the whole system. And the statistics from this MLH1 screen are profoundly validating for the platform.
14:20I mean, when they analyze variants that clinical databases already confirmed, we're pathogenic. 54% of them scored clearly as loss of function in the prime editing screen. Okay, 54%. Conversely, only 2.4% of known benign variants behave that way.
14:34That is a huge difference. That clear separation means the platform is highly accurate at discerning the dangerous typos from the harmless quirks. And they managed to reclassify 7.5% of the mysterious unclassified variants tested just one specific section of the gene as loss of function.
14:53They literally cleared away some of the fog in the global clinical database in a single sweep. Yeah, and the primary clinical relevance here really comes down to biological context and scalability, because previously, testing an unclassified variant often meant extracting the gene, putting it into an artificial loop of DNA called a plasmid and forcing a cell to overproduce it.
15:14Which isn't how it works in our bodies. Exactly. The protein levels were entirely unnatural. This platform proves we can scalably assess variants right in their native genomic context, keeping the natural biological balance intact.
15:25Which is the key to finally clearing that massive backlog of mysteries in clinical databases. If evasion is told they have an unclassified mutation. We now have a viable, high throughput pathway to functionally test that exact mutation and give them a definitive answer.
15:42It's a game changer for diagnostics. But the ripple effects go beyond the Petri dish. I'm looking at how this feeds the computational biology world, too. Oh, totally. The AI models, like splice AI that try to predict variant effects computationally, well, they rely entirely on their training data.
15:58You cannot train an algorithm to recognize the incredibly subtle multidimensional rules of human genetics without 100s of 1000s of highly accurate real-world data points. These massive data sets generated by prime editing screens, especially in those tricky non-coding regions, are the exact caliber of data those predictive models desperately need to improve.
16:20Okay, let's talk about the limitations, though, because no scientific breakthrough is without its caveats. I was looking at their quality control metrics, specifically that surrogate target filter we discussed earlier.
16:30Right, the GPS tracker. Yeah. During the Sam Marci B1 screen, they demanded a PEG RNA show greater than 75% editing efficiency on the surrogate target before they would trust its biological result. Wait, 75%.
16:44If you toss out everything below that, aren't we creating a massive blind spot, like how many potentially fatal disease variants are we just ignoring because the guide RNA was slightly inefficient? Yeah, that is a very real, very transparent limitation of the current study.
17:01The high threshold guarantees that the variants passing the filter are accurately assessed, preventing false positives, which is absolutely vital. If this data is going to inform clinical decisions. But you are entirely correct, it throws out a massive amount of data.
17:14It severely reduces the sheer number of variants they can confidently score in one single experiment. It's a calculated and frankly somewhat painful trade-off between the quantity of data and the quality of data.
17:25So we have a solid framework, but the prime editing engine itself needs more horsepower. How do they solve that massive data loss moving forward? The immediate next steps involve upgrading the prime editing machinery itself?
17:38The authors note that adopting newer generations of prime editors like PE7, which have been engineered for much higher baseline efficiency, will help immensely. Ooh, P7, okay. They're also pairing this with improved algorithmic design for the Pedge RNAs themselves.
17:53By using machine learning to predict which guide RNAs will be highly efficient before anyone even synthesizes them in the lab, they can drastically lower the dud rate. And by boosting that overall efficiency, they mention the ultimate goal is to scale this up to test all 2000 plus essential genes in these haploid cells.
18:12The group of that is breathtaking. It's like proofreading the most critical chapters of the human instruction manual, one single letter at a time. It is a monumental task, but one that is finally within our technical reach.
18:23We are transitioning from a passive era of just reading genetics to an active functional understanding of our own biology. If we distill this deep dive down, Pooled prime editing, enhanced by ingenious quality controls like surrogate targets, allow scientists to test 1000s of genetic variants directly in their natural genomic context, even in the vast non-coding regions.
18:44This platform is a massive practical leap toward turning the unknown dark matter of our DNA into actionable clinical knowledge, which brings up a fascinating thought to mull over. If we are currently mapping out exactly which single letter typos break human health, and we have the exact precision prime editing tools to prove it in the lab, how long until we aren't just using these tools to diagnose the typos, but deploying them as therapeutics to actively rewrite those errors in living patients?
19:12That is the ultimate horizon we are all looking toward. What does this mean for the 1000000s of people carrying rare unclassified genetic mutations? It means we are finally building the dictionary they need to understand their own DNA.
19:24This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
19:39If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
19:49Thanks for listening and join us next time as we explore more science, base by base.