0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Glad to be here for another deep dive.
0:10Yeah, so imagine for a 2nd that you are participating in this highly controlled, honestly incredibly bizarre experiment. Okay. I'm imagining it Right. So you are standing in a hallway, and there are 3 completely separate locked door.
0:27Locked doors, got it. Yep. And behind each door is a room, and inside each room, you are sitting there, meticulously copying a completely different chapter of a book by hand. Just writing it all out Exactly.
0:39So, roommates chapter one, roomist chapter two, room C is chapter three. Right. You finish your work, you lock the doors behind you, and you slide the pages under the doors into this, like, secure collection box.
0:51So they're totally isolated. Completely. There is absolutely 0 physical way for the pages to interact with each other. Okay, follow you. But when you go to read the copies you just collected from the box, something logically impossible has happened.
1:04Yeah. The chapters have somehow inexplicably mashed themselves together into a single seamless Frankenstein paragraph. Wow. Right. Like a sentence from chapter one flows directly into a sentence for chapter 3 on the exact same continuous piece of paper.
1:21That makes no sense. Exactly. How could this happen? What really happens when the very code of life gets inexplicably scrambled in a locked room? Well, welcome to this deep dive. It really is the ultimate locked room mystery, but, you know, playing out on a molecular level.
1:37It really is. And it's one that has massive implications for anyone relying on genomic data today. Absolutely. And today we celebrate the work of a brilliant research team. Ruby White, Christoph Pelfigs, Olivier Lomia Bell, and David Eccles.
1:51A really fantastic group. Yeah, they published this fascinating 2017 research paper in F 1000 research, and it was titled Investigation of Chimeric Reads Using the Minion. It's a classic in the field at this point.
2:03It really completely turns this locked room mystery on its head. So okay, let's unpack this. Let's do it. Our mission for this deep dive is all about what happens when genetic sequencing goes weirdly off script.
2:14We're looking at the hidden, highly structured errors that can, you know, slip past our technology and actually masquerade as real biology. And we need to establish right away that this isn't just a quirky lab accident.
2:27Right. It's not a mistake Exactly. It's not some clumsy technician mixing up tubes on a bench. This is a fundamental discovery about the mechanics of how we actually prepare and read DNA. The chemistry itself.
2:39Yeah, the chemistry and the artificial computational illusions we can accidentally create in the lab. Okay, so we find our researchers in the lab. And they're conducting a nanopore sequencing run. Standard stuff.
2:52Right. They are sequencing PCR products from 3 different amplicons. And just to set the scale, all of these are less than one kilo base in size. So pretty small fragments. Yeah. And we know our audience is familiar with standard PCR amplification, so we don't need to re-explain the whole molecular photocopier concept.
3:10We can skip that. What we do need to talk about is the specific technology they were using because while they immediately ran into a massive problem. A really big one. Yeah, an unexpected abundance of their sequencing reads, completely failed quality control.
3:24The machine just rejected them. Exactly. The machine started tossing the data out into the digital trash. Just failing rate after red. Yeah, and the specific reasons cited for this failure by the software was a template complement mismatch.
3:39Which is a very specific type of error. Right. And I want to make sure I'm visualizing this correctly. Wait, a template compliment mismatch. What does that actually mean in the machine? Like, why does it cause the whole read to fail quality control?
3:51It's interesting because. Is it like a zipper where the teeth suddenly don't align and the physical thing just gets stuck? That is actually a much more accurate way to look at it than assuming a physical jam in the machine.
4:02Oh really? Yeah. It is a computational, logical failure, not a physical one. The poor itself was pulling the DNA molecule through just fine. So it wasn't stuck. No, not at all. You see, the machine reads the 1st side.
4:17Uh, the template strand. Okay. Then the sequencing chemistry utilizes a hairpin curve. This allows the machine to seamlessly turn the corner and read the corresponding opposite strand. The complement strand. Exactly, the complement.
4:32And the software is constantly checking its work as it goes. proofreading. Right. It reads the 1st half, and it expects the 2nd half to be a perfect biological mirror image. Because A pairs with T and C pairs with G.
4:45Precisely. But in this specific run, the software was reading the template strand, and when it turned that hairpin corner to verify against the compliment strand, it encountered a totally different sequence.
4:57Just totally non-matching. Right. The logical pair was completely broken. Oh wow. The machine effectively said, you know, this is making biological sense. The English half doesn't match the Spanish half.
5:07The data is corrupted. So it flagged it as a QC failure. Yep. throw in the reed right into the reject pile. Man. A massive QC failure like that demands a serious investigation. I mean, you don't just throw away a huge chunk of your data without asking why the logic broke.
5:24You can't. You have to know why. So the researchers put on their molecular detective hats. They ran a blast search on these failed rejected reeds. Which is a standard way to map sequences to known databases.
5:38Right. And here's where it gets really interesting. Oh, this is the best part. The blast search revealed that some of these failed single continuous reads match perfectly to 2 entirely different amplicons.
5:50The 2 different genes. Yes. A single intact piece of DNA was half gene A and half gene B. Which shouldn't be possible. Exactly, because the paper explicitly points out that the PCR amplification was carried out completely separately for each amplicon.
6:05Right, they were in separate tubes. They were in separate tubes. So if they amplified an entirely different environments, how on earth do 2 separate genes physically fused together into a single molecule?
6:16It completely defies the basic physics of the experiment as it was initially set up. It's the lock room. Yes. You have Amplicon A, multiplying in 2A, and Amplicon B multiplying in 2B. You eventually run them through the sequencer, and out comes a single continuous piece of DNA that the blast algorithm recognizes as half A and half B.
6:38The paper calls these chimeric reads named after the chimera from Greek mythology. Frankenstein paragraph from our locked room analogy. Right. Now, as a rigorous scientist, you have to rule out human error first.
6:50Of course. You assume you messed up. Exactly. Did someone accidentally cross-contaminate the tubes before the sequencing prep? Did someone use a dirty prey pet? Right. So the team had to design a very specific intervention to prove these chimeras were a real chemical phenomenon and not just sloppy pipe heading.
7:09Okay, so they isolated the variables. To specifically hunt for these chimeric reeds, they use completely separate barcodes for each amplecon. Crucial step. And they tested 2 different legation methods prior to loading the samples onto the minion.
7:22Let's dig into the mechanics of that. I know barcodes are used as molecular ID tags, but if these sequences are physically fusing into chimeras, aren't the barcode tags getting scrambled too? You would think so, but... Like, how did they untangle that bioinformatically?
7:38What's fascinating here? Is the sheer elegance of how they use those barcodes to map the chaos? Okay, lay it on me. So molecular barcode is basically attached to the ends of the amplicon. Like a shipping level?
7:51Exactly like a shipping label. By attaching a distinct, unique barcode to Amplicon A, a different one to Amplicon B, and a 3rd to C, you can track exactly where every part originated. Okay I'm with you.
8:05If you find a single continuous read in your data that possesses the barcode for Amplicon A on the far left end. But the genetic sequence of Amplicant C in the middle, and maybe barcode C on the far right end, you have definitive proof.
8:18Oh I see. You have untangleable proof that a physical fusion happened. The tags didn't scramble. They just framed the fusion. Yeah, exactly. They framed a physically fused chimeric molecule. Wow. Which brings us to the ligation methods.
8:31Right, the glue. Yeah, if they amplified in separate tubes. The only time these amplicons physically meet is when they are finally pulled together to be prepped for the sequencer. That's the only time they are in the same room, so to speak.
8:43And that involves ligation, the molecular glue. Walk us through what is physically happening in that chemical bath that could cause this. Okay, think about the physical reality of a ligation bath. You are taking your 1000000000s of copies of Amplicon, A, B, and C and you are pulling them into a single microcentrifuge tube.
9:02So it's very crowded in there. Extremely crowded. You then add your sequencing adapters and you add LIGUS. The enzyme. Right. An enzyme whose sole job is to glue pieces of DNA together. Okay. The intended desired reaction is that the Lygus glues an adapter to the end of an amplicon.
9:20That's what prepares it for the sequencing pore. But Legas isn't smart. Right? At all. Like it doesn't know, it's only supposed to glue adapters to amplicons. It's just a dumb chemical catalyst. Exactly.
9:30It has no brain. It's just a numbers game based on concentration and thermodynamics. Okay, so what happens? In that crowded microscopic environment? An Amplicon A molecule is bumping into adapters, sure, but it's also bumping into Amplicon B molecules.
9:46Oh, I see where this is going. Yeah, sometimes the Legus enzyme grabs Amplicon A and Amplicon B and just glues them end to end. Just accidentally. Totally accidentally. It forms a highly stable, artificial chimeric molecule before the sequencing adapter even gets attached.
10:01So after systematically isolating these variables, keeping the implicants separated during PCR, using strict barcoding and managing the legation products, they look at the base called sequence, and the chimeric reads were still there.
10:14They didn't go away. They survived the prep. The text gives us a very specific and frankly startling piece of hard data. numbers are really wild here. Yeah, it states that at least one. 7% of the reads prepared using the nanopore LSK 002 2D ligation kit included these post-amplification chimeric elements.
10:33If we connect this to the bigger picture, That one. 7% is a staggering statistic for anyone working in genomics. I have to push back on that a bit, though. Okay why? To someone outside bioinformatics, one.
10:477% sounds like a tiny fraction. It does sound small. Right. It's less than 2 in a 100. In a sequencing run of 1000000s of reads, wouldn't standard error correction algorithms just like filter that out as statistical noise?
11:00You hope so. So why is this such a dangerous blind spot? It's a brilliant question, and it really gets to the heart of why this paper is so important. Okay. You were right that algorithms filter out noise.
11:11Like random single letter sequencing errors. They're just static. Right. static. But a chimeric read doesn't look like static or noise to the software. What does it look like? It looks like a highly structured, perfectly valid biological event.
11:25Oh, no. In genomics. real biological fusions happen all the time. Think about cancer research. We actively look for translocation events where 2 separate genes fused together to cause a tumor. Oh, wow.
11:39So if a cancer researcher sees a continuous sequence containing half of gene A and half a gene B, they don't think error. Absolutely not They think they just found the driver mutation for a new cancer type.
11:50Precisely. If 1... 7% of a million reads are chimeric. That means you have 17,000 Frankenstein molecules in your data pool. That is terrifying. 17,000 highly convincing illusions suggesting 2 genes are naturally fused together.
12:06When, in reality, they were just accidentally glued together in your sample preparation tube by an overactive legus enzyme. Just a thermodynamic accident. Exactly. If you don't realize those are artificial, it could completely skew your understanding of the biological system you're studying.
12:21The algorithms keep them because they look biologically real. So who is the bad guy here? That's the big question. Is this just a specific flaw in that nanopore LSK 0022D legation kit? Like, if I'm a researcher using a different platform, Am I safe?
12:37This raises an important point, and the paper is very careful to directly address it. Okay, what do they say? While this specific 1... 7% rate was observed using that specific nanopore kit, the researchers concluded that this chimera forming process is unlikely to be specific only to the sample preparation used for nanopore.
12:57Unlikely to be specific only to nanopor. Okay. Because the concept of ligation, you know, of putting amplicons into a crowded chemical bath with lycus enzymes to attach adapters, that is a universal vulnerability.
13:11Pretty much all high throughput sequencing technologies rely on a ligation step to prep the DNA library, regardless of the brand of the sequencer. Exactly. The laws of chemistry don't change just because you bought a different, more expensive sequencer.
13:26Whenever you have lots of highly concentrated DNA fragments and molecular glue in the same environment, you're going to get artificial fusions. It is a fundamental thermodynamic risk of the prep phase itself.
13:37But here's where the paper offers a massive silver lining, and it actually vindicates the specific technology they were using. Great twist. Yeah. So what does this all mean? The paper points out that the long read nature of nanopore sequencing presents an incredibly effective tool for the discovery and filtering of these chimeric reads.
13:56This is the crucial twist, yeah. But wait, we just said these fusions look like real biology? How does a long read technology actually help us spot the illusion? It comes down to context. context. Yeah, think back to our Frankenstein paragraph or even our translated document analogy.
14:12If you are using an older, traditional short read sequencing technology. The machine only reads tiny, fragmented snippets of DNA at a time. Like how tiny? Maybe 150 base pairs. So it's looking at the book through a tiny keyhole.
14:26Yes. A very small keyhole. And if that keyhole happens to land exactly on the junction where Amplicant A was glued to Amplicon B. I see it. The short rate sequencer just outputs a tiny snippet that is 75 bases of A and 75 bases of B.
14:40And to the researcher, it looks like a perfect flawless biological fusion. Exactly. There is 0 broader context to tell them it's an artifact. Because it just reads the joint. It doesn't see that the rest of the gene isn't actually attached correctly.
14:53Exactly. But nanapour is a long read technology. It pulls massive, continuous strands of DNA through the poor. It reads the whole page. or reads the whole page, not just the keyhole, because it reads the whole lengths.
15:06A researcher or, you know, a bio-informatics pipeline. can look at the data and see the entire structure. They see the big picture. Right. They see the barcode, the entirety of Amplicon A, the unnatural joint, the entirety of Amplicon B, and the other barcode.
15:21Wow. The long read nature provides the full sprawling context of the error. It makes the unnatural junction glaringly obvious in the wider landscape of the sequence. And that allows developers to write software that specifically flags and filters it out as an artifact before it pollutes the downstream analysis.
15:41So the very technology that was used when they discovered this one. 7% chimera rate is actually uniquely equipped to identify and computationally discard them. Exactly. It's the cure as well as the diagnostic.
15:53If we synthesize the core takeaway from this incredibly elegant piece of scientific detective work. It really serves as a massive wake-up call for the entire field of genomics. It really is The researchers prove that post amplification chimeras are a real, measurable, and highly disruptive phenomenon.
16:11And crucially, the warning extends far beyond just this one experiment or this one specific 2D kit. It's an industry wide warning. Right. It implies that any amplicon sequencing process that relies on a pooling and ligation step is actively at risk of generating these artifutal fusions.
16:29Wow. The global scientific community has to be acutely aware that the very chemistry we use to prepare our samples can fundamentally alter the biological truth of the data we receive. It's a brilliant reminder that our tools, you know, no matter how advanced are not magic boxes.
16:45They are definitely not magic. They are subject to the messy realities of chemistry, and the data they output must always be subjected to rigorous, highly skeptical analysis. We have to thoroughly understand the thermodynamic quirks of the sample prep, to truly trust the biology we think we are discovering.
17:02Because without that rigorous skepticism, we run the risk of cataloguing a biological world that only exists inside a microcentrifuge tube. Which brings up a final, somewhat haunting thought for you to mull over as we wrap up this deep dive.
17:17Oh, this a good one. We've established that these chimeric elements are happening post-amplification entirely due to standard ubiquitous chemical prep steps. We know they occur at a highly significant rate, at least one.
17:31in this specific study. A huge number. And we just discussed how traditional short red technologies are basically blind to them, viewing the unnatural joints through a keyhole and interpreting them as real fusions.
17:43They can't see the context. So think about the data. Think about the massive global databases of genetic information sequenced over the last 15 years, heavily reliant on older, short red technologies that couldn't easily spot these whole page fusions.
17:58It staggering to think about. How much historical genetic data out there might unknowingly contain these tiny artificial genetic mashups. How many novel transcripts or fusion genes in our public databases are just quiet, unrecognized chemical artifacts hiding in plain sight.
18:16It is a profound, and honestly, somewhat unsettling question. And one that perfectly highlights why the continuous refinement of our bioinformatics pipelines and studies like this one are absolutely critical to the integrity of modern science.
18:30This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
18:43If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
18:52Thanks for listening and join us next time as we explore more science. base by base.