This episode reviews a Clinical Chemistry study that developed an improved one-step ultra-long-range PCR with PCR suppression primers and PacBio SMRT long-read sequencing to obtain 26.1 kb full-length ABO haplotypes from the 5′ UTR to the 3′ UTR, enabling comprehensive allele annotation and resolution of complex ABO variants.
0:09On bright screens, the letters fall in line, free at a time, like a clockwork sign, but in this quiet cell room. Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app.
0:27Imagine you are a clinician, right? You're standing in the emergency room. Oh, yeah, that is a terrifying scenario from the jump. Exactly. A patient desperately needs a transfusion, so you run a routine blood draw.
0:37But the lab results that come back are, well, biologically impossible. Like you're looking at a sample that somehow has a conflicting mix of entirely different cell types, just actively fighting a turf war in the vial.
0:51Or maybe the patient tests as this incredibly rare, weak variant of type B. Which barely even registers on the hospital's diagnostic machines. I mean, a discrepancy on a blood test isn't just some academic puzzle you can set aside for peer review.
1:05It is a life or death crisis happening right there in the room. Our standard labels, you know, A, B, A, B, and O, they are universally understood. But when they fail that baseline expectation of total medical certainty just evaporates.
1:17So what really happens when those standard labels break down. And how do we solve these life-threatening blood typing discrepancies? To understand this, we have to recognize that our traditional way of blood typing and even our standard genetic sequencing has been like reading a heavily redacted document.
1:34Exactly. We are finding out that the solutions to these massive medical mysteries are actually hidden deep in the unedited non-coding margins of our DNA. How could this change the way we match blood for life saving transfusions?
1:48Well, to answer that, we really have to look at a massive technological leap forward and how we actually read that DNA, because for a long time, the tools we relied on were inherently limiting our perspective on the human genome.
2:01Today we celebrate the work of the team at the blood center of Zejing Province, and the blood transfusion Medicine Research Institute in Hangzoo, China, who have advanced our understanding of blood group genomics by looking at the entire ABO gene from end to end.
2:14Yeah, and that end-to-end aspect represents a total paradigm shift. Historically, even with the incredibly sophisticated, high throughput sequencing technologies we've developed. We have only been looking at a fraction of the genetic picture.
2:28Pass sequencing efforts focused almost exclusively on the coding DNA sequences. The Exxons. Right, the Exxons. Which are the specific instructions for building proteins. Like in the context of blood types.
2:40We are talking about glycosol transfraces, right? The enzymes that do the physical labor of attaching the A or B sugar antigens to the outer membrane of your red blood cells. Right, but a gene is so much more than just its excellence.
2:52The ABO gene possesses a vast amount of genetic dark matter. This includes introns, the intervening sequences between the exxons, as well as the 5 prime and 3 prime untranslated regions, or UTRs. Those are the parts flanking the gene, the very beginning and the very end.
3:07Exactly. And for decades, these non-coding regions were dismissed by early geneticists as junk DNA. Which goes back to our redacted document metaphor. If the exxons are the action verbs of the sentence, the words telling the cell to run or jump.
3:22Previous sequencing methods only let us read those isolated words. We were completely missing the grammar, the punctuation, and the structural formatting that gave those words their actual meaning. And that punctuation dictates everything.
3:35These non-coding UTRs and introns are packed with regulatory elements. They serve as docking stations for transcription factors. They dictate MRNA stability, and they tell the cell precisely when and how much of that glycosol transferase enzyme to produce.
3:51So if a mutation occurs in that dark matter, it drastically alters the enzyme's activity, which is exactly what creates those rare, mysterious blood variants that stump standard hospital tests. I'm stuck on something, though.
4:04What's that? If we know the UTRs and introns are doing all this heavy lifting, it feels like malpractice that we haven't just sequenced them from the start. What has actually been starping us from reading the unredacted document this whole time?
4:15It comes down to sheer physical distance and structural chaos. The ABO gene is remarkably large. The distance from the 5 Prime UTR at the start of the gene to the single nucleotide variants that determine your blood type, which are way down in Exxon 6 and 7, is immense.
4:31Wow, okay. And to complicate matters, the 3 prime UTR at the very end of the gene is notoriously chaotic, is full of highly variable, short, repetitive sequences. Like a record skipping over and over in the exact same groove.
4:44That's a perfect way to put it. When traditional sequencing technologies try to read those highly repetitive areas, they inevitably lose their place. Because of this blind spot, even the standard global reference sequences we use in medicine have lacked complete 5 Prime and 3 Prime UTR data.
5:00Conventional long range PCR, the process we use to copy DNA to study it. Well, it starts to fail when trying to amplify DNA fragments that are over 10 kilobases long. And the complete ADO gene is significantly larger than 10 kilobases.
5:13It is. So the researchers in Zhijing had to engineer a workaround. They optimize a technique called one-step ultra long range PCR, or ULR, PCR. The critical innovation here was the implementation of a concept called PCR suppression, utilizing a highly specialized pair of primers.
5:32Let's break down that mechanism. I know PCR relies on primers, which act as short starting blocks of DNA that signaled the replication enzymes where to begin copying. How does a primer suppress the process to capture massive fragments instead of the usual short ones?
5:47So they designed these primers with a very specific 25 nucleotide sequence appended to the 5 Prime end. This sequence was designed to be extremely GC rich. Because guanine and cytocene bonds are incredibly sticky and stable.
6:00Highly stable. During the PCR thermal cycling process, You have an abundance of short, single stranded pieces of DNA floating around. Because those added ends are so sticky, the short DNA strands fall back on themselves.
6:11Ooh, weird. Yeah, the ends snap together and lock, forming tight loops that structural biologists call panhandle like structures. Literally resembling the handle of a frying pan. Exactly. The sequence binds to itself and physically locks up because the ends of those short strands are locked in this panhandle shape.
6:28The primers can no longer bind to them. The replication enzyme simply cannot access the starting line. This physical block changes the entire landscape of the reaction. So the researchers essentially trick the assay.
6:40By putting a chemical padlock on all this short, easy to read fragments, they force the entire replication machinery to focus exclusively on copying the massive, difficult, ultra long fragments. It completely shifts the amplification bias.
6:55By suppressing the short fragments. They successfully amplified an unbroken stretch of DNA that was 26.1 kilobases long. That's massive. It's huge, but just copying a massive fragment isn't enough. You know, you still have to read it.
7:08So they fed those ultra long amplicons into a single molecule real-time sequencing platform, specifically the packed bio sequel 2. How does that platform differ from the standard sequencing we usually see where everything is chopped up into tiny reeds and then reassembled like a jigsaw puzzle by a computer?
7:24Well, single molecule, real time, or SMRK sequencing is a marvel of biophysics. It doesn't chop anything up at all. Instead, it uses microscopic wells called 0 mode wave guides. These wells are so incredibly tiny that light can only illuminate the very bottom of them.
7:41At the bottom of each well. They anchor a single polymerase enzyme. Oh wow. Yeah. As that enzyme pulls the massive DNA strand through and attaches fluorescently tagged nucleotides to build the complementary strand, it emits a microscopic flash of colored light.
7:57We are literally watching a single molecule of DNA being synthesized in real time base by base, completely uninterrupted. They force the chemistry to hand them the entire uncut manuscript and then they watch it be transcribed live.
8:10Exactly. Now that they built a tool capable of reading the whole unbroken sequence, they immediately aimed it at the very mysteries that started this conversation. Those impossible blood types in the emergency room.
8:20They tested this on 79 random healthy donors and 47 patients with known ABO variants. What did that massive 26. view actually reveal? It provided the most complete ABO gene map ever reported in human history.
8:33No splicing, no algorithmic guesswork. First, they mapped out the normal donors, classifying the 5 predominant ABO alleals in the Chinese population. That's A1.01, A1.02, B.01, 0.01.01 and 0.01.02. Right.
8:50But because they could finally read the dark matter, they subdivided these common alleges based on newly discovered structural architecture deep in the introns and regulatory regions. So they found hidden variations within blood types we thought we completely understood.
9:02They identified variable tandem repeats, or VNTRs, inside Intron one, specifically repeats of the base's TNA. They found that the A1.02 and B .01 alleles contain an average of 21 of these TA repeats. The OELs, however, only average about 13 repeats.
9:19But why do those repeats actually matter? Do a few extra stutters in a non-coding region fundamentally change how the gene operates? It completely alters the structural scaffolding of the gene itself. When you add or remove tandem repeats in an intron, you change the physical spacing of the regulatory elements.
9:35It can affect how the DNA bends or how the resulting RNA transcript folds onto itself. The physical architecture dictates the biological function. Which proves the junk DNA theory was completely blind to structural mechanics.
9:48Oh, totally. And it didn't stop there either. In the 5 prime ETR, they mapped out specific 43 base pair mini satellite repeats. And in that notoriously chaotic skipping record region of the 3 Prime UTR, they categorize the sequence into 14 distinct structural units.
10:0514 units. Earlier we noted that the standard global reference sequence lacked data here. How did this new 14 unit map compared to the global reference that the medical community relies on? This is where the data becomes deeply consequential for global medicine, the standard global reference sequence, NG 00669.2, is definitively missing units 11, 12, and 13 compared to the Chinese population data.
10:28Wait, seriously? Seriously. The foundational map we've been using to understand human blood is literally missing paragraphs at the end of the document. That is a staggering oversight in our baseline reference material, but let's look at the immediate clinical impact.
10:42What happened when they applied this ultra long red to the 47 variant cases. The patients whose blood typing was causing those life or death discrepancies. The continuous sequence resolved them beautifully.
10:53They uncovered three entirely novel coding variants. More remarkably, and one patient presenting with a massive typing discrepancy, they found a deletion of 7337 base pairs, just massive canyon of missing genetic material.
11:08Over 7000 base pairs deleted. How is that person's blood even functioning? Yeah, the key is where the deletion occurred? It didn't destroy the coding sequence itself. The Exxons were completely untouched.
11:19Instead, that 7000 base paradeletion completely erased a specific transcription enhancer region located far upstream in an intron. Walk me through the mechanics of that. If the Exxons are intact, the blueprint to build the blood antigen is there.
11:32How does an enhancer missing 1000s of base pairs away silence the gene. We have to think about the genome in three dimensions. DNA isn't a stiff straight line. You know, it loops and bends. An enhancer acts as a critical docking station for proteins that rev up the genes activity.
11:50Okay. In a normal cell, the DNA physically loops over so that the distant enhancer touches the promoter region at the start of the gene, turning the volume up to 10. If you delete that docking station, the structural blueprint for the blood type remains intact, but the gene is permanently dialed down to a volume of one.
12:08The instructions are there, but the cell just barely whispers them. Precisely. This patient had what is known as a Bell phenotype. Their blood cells technically had the B antigen, but because that enhancer was missing, the expression was incredibly weak.
12:22Standard sequencing, which only looks at the Exxons, would see a perfectly normal type B blueprint, and misdiagnose the severity of the issue, leaving clinicians confused when the actual blood sample barely reacts.
12:33That makes so much sense. They also mapped recombination hotspots, areas where alleles physically swapped genetic material during cell division, creating hybrid phenotypes like ABW. There was one specific clinical mystery they solve that completely shifts how we view human biology.
12:49They analyzed a case of mixed field surology, where the patient turned out to be a microchimera. We hinted at this in the beginning, a turf war in the blood vial. Can you explain the biological reality of how someone can have multiple blood haplotypes circulating in their body simultaneously?
13:07Yeah, so a chimera is a single organism composed of cells with more than one distinct genotype. While it can happen after a bone marrow transplant, it can also happen naturally in utero. If 2 fraternal twin embryos fuse very early in development, the resulting person is born with 2 completely different sets of hematopoietic stem cells.
13:25Wait, so their bone marrow is literally a patchwork of 2 distinct genetic identities. Yes. And they are biologically their own twin, and their bone marrow is pumping out competing blood types. That is wild.
13:36It is. And when this patient's blood was tested in the hospital. The serology showed mixed field agglutination. A small minority of the cells reacted to the A antigen test, while the vast majority didn't.
13:49Naturally, the hospital ran a standard PCR sequencing test to figure out what was wrong. But standard PCR chops the DNA into tiny fragments before reading them. The irony is that the standard test actually did exactly what it was programmed to do, which is why it failed so spectacularly.
14:06When it chopped up the Chimera's DNA, it just saw a massive chaotic alphabet soup. It found markers for type B, markers for typo, and a tiny trace of type A. I see. Because it couldn't tell which markers belong to which cells.
14:20The diagnostic software assumed it was looking at one incredibly weird, mutated hybrid allele, and misdiagnosed the patient as a simple header as I goat. Right, because if you take a redacted manuscript, shred the pages, throw them in a pile, and hand them to an algorithm.
14:34Can't tell if it's reading one story told from 3 perspectives or 3 completely different books tossed into a blender. Exactly. But because this ULR PCR method reads the entire 26.one kilobase fragment unbroken, it doesn't just read the genetic letters, it reads the phasing, it reads the complete sentences.
14:52It definitively proved that the patient's blood wasn't one mutated illille. It was a true biological mixture of 3 distinct cell lines. Roughly 12% of their cells were operating entirely on the A1.02 haplotype instruction manual, coexisting alongside completely separate populations of B .01 and 0.01.02 cells.
15:13The unbroken string proved that 12% of this person's blood was operating on a completely different genetic reality. Seeing how this unredacted sequencing resolve, massive enhancer deletions and hidden chimerism, really begs the broader question.
15:25If we connect this to the bigger picture of clinical practice, how will science apply this 26 kilobays view moving forward? Well, the implications for both diagnostics and basic biology are vast. First, capturing the complete 3 prime UTR sequence finally allows us to study MRNA stability in blood phenotypes.
15:41The 3 Prime UTR accesses or countdown timer for the MRNA, right? Yeah, that's a perfect way to conceptualize it. Jeans aren't just flipped on or off. The Messenger RNA that carries the instructions has a definitive lifespan.
15:52The 3 prime UTR, folds into physical stem loop structures. Cellular enzymes slowly chew away at these loops over time. Once the loops are destroyed, the MRNA degrades. This structural timer dictates exactly how long the instructions to build blood antigens stick around in the cell.
16:10If a variant alters those loops, the timer might run out too fast, leading to a weak blood expression. Furthermore, reading the unbroken string allows us to map single nucleotide variants, the SNVs deep within the introns, which acts as an incredible new tracking system.
16:26We already know that certain ABO blood types are linked to susceptibility for various conditions, including pancreatic cancer and cardiovascular disease. By mapping these newly discovered SNVs across the whole Hapla type, researchers can trace these specific structural variations to better understand the underlying mechanisms behind those disease associations.
16:44The researchers actually suggest this technology could lead to a multilevel naming system for ABO alleles, kind of similar to how we categorize human leucocyte antigens, HLA for organ and stem cell transplants.
16:57You wouldn't just be type A. You would have a highly specific numerical code describing the exact regulatory and structural makeup of your blood. It introduces an unprecedented level of precision to personalized medicine.
17:09It is vastly superior to next generation sequencing for detecting large structural variations. However, any new diagnostic frontier comes with significant physical limitation. Right. What is the blind spot for this ultra long amplification?
17:24Biology is fragile. To amplify a continuous 26.1 kilobase fragment, you require highly intact, high quality DNA. If the DNA in the sample is fragmented or degraded, the panhandle suppression technique cannot magically piece it back together.
17:38There is a deep irony there. We have engineered this almost magical sequencing technique, capable of reading 26,000 base pairs in real time, yet its success is completely at the mercy of whether the clinical staff stored the biological sample at the correct temperature in the fridge.
17:53Oh completely. It is a stark reminder of the physical realities of lab work. Additionally, because these variants were discovered so deep within newly mapped introns, we currently lack established secondary methods to easily verify them.
18:07Standard clinical tools simply cannot reach the genomic depths this assay just explored. Consequently, we are heavily reliant on the internal accuracy of the single long read approach until secondary verification tools catch up.
18:20But even with those physical limitations, this represents a monumental leap forward in clinical diagnostics, to synthesize everything we've unpacked in this deep dive into the research, full-length, continuous ABO haplotype sequencing from the 5 Prime UTR straight through to the 3 Prime UTR has finally illuminated the hidden genetic architecture of our blood.
18:41It really has. It has given us the tools to read the regulatory dark matter that the medical field has been forced to ignore for decades. And this unbroken 26.1 kilobase perspective allows clinicians to resolve immense structural variations and uncover mysterious blood phenotypes, like true microchimerism that older segmented methods were structurally blind to.
19:02So, I will leave you with this final thought. What does this mean for the future of personalized medicine when mapping a single gene cover to cover can reveal hidden cellular chimeras and rewrite how we understand our own blood.
19:13It suggests we are finally moving past the illusion that we are just the simple sum of our A, B, and O labels. We are the complex product of the silent, unedited margins of our DNA, and we have only just begun to learn how to read the complete manuscript.
19:26This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating.
19:40If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
19:50Thanks for listening, and join us next time as we explore more science base by base. On bright screens, the letters fall in line free at a time like a clockwork sign. But in this quiet cell, the rules feel new to stoplight shimmer into something true.
20:20You a won't let the story die. You, AG, won't say goodbye. A tiny adapter changes the view and the sentence keeps running through. Stops, I turned the lyrics in the code tonight, UAA delights. You ate the glue, all right, red fool, like a river when the gate comes loose.
20:46But you gay stands guard like a final truce. Found the TRNA shape just right and take a donkeys in the lab low light. One points to Licey. One to glue to made, rewriting endings at the rabble some's gate.
21:11Still downstream. There's a double stop line, tandem UGA, but design, a safety net. Where the last word land, so proteins don't spill past. The plan stops. I turn to lyrics in the code tonight, UAA delights, you ate to glue, all right, annotation, dreams need a wider lens.
21:37'Cause the code can change where life begins.