This episode reviews a Cell Genomics study that uses ssDNA '2-end' and novel '4-end' sequencing to profile native 5′ and 3′ termini of plasma cfDNA. The work identifies PREM/POEM markers, links 3′ ends to methylation, and shows improved HCC detection.
0:00Welcome to Base by Base, the papercast that brings ginomits to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So, um, I want you to imagine trying to read an ancient priceless manuscript, you know, something that holds the secrets to a lost civilization.
0:18Oh, like a super fragile, crumbling document. Yeah, exactly. But to fit this delicate document into a modern digital scanner. Someone has basically taken a pair of scissors and just sanded down all the torn, jagged edges of the pages.
0:33Wait, just to make them perfectly straight, that's horrifying. Right. I mean, you'd still have the text in the middle, sure, but you would lose critical clues about how the book was torn, or, you know, what kind of damage it suffered, and really what happened to it over the centuries.
0:46Which would be a total nightmare for a historian. It really would. But the crazy thing is, this is exactly what we do right now in standard liquid biopsies for cancer. Yeah, it's um, it's kind of an open secret in the field.
0:58Right. Because we look at fragments of tumor DNA floating in the blood, what we call self-free DNA or CF DNA. But our standard chemical processing actually destroys the natural ends of those DNA fragments.
1:11It literally chops away the evidence. Exactly. So what really happens when we stop destroying this evidence and, you know, look at the whole picture. How could reading the raw, untouched edges of our DNA change our ability to detect cancer before a tumor is ever seen on a scan?
1:28I mean, that is the multimillion dollar question in fragmentomics right now? Because for a long time, we just sort of accepted that losing those edges was the cost of doing business. But questioning that assumption is fundamentally changing things, right?
1:40Oh, absolutely. It is completely changing how we view these floating bits of genetic information. What today we celebrate the work of Peong Jiang, YM Dennis Lowe, and their incredible team at the Center for Navostics and the Chinese University of Hong Kong, who have really advanced our understanding of self-free DNA fragment comics.
1:58Yeah, their work is just phenomenal. It really is. We are doing a deep dive into their groundbreaking work, published in cell genomics on March 11, 2026. And, you know, to really grasp the magnitude of what they've achieved here.
2:11We should probably take a step back and look at the mechanical problem that has been plaguing CFDNA analysis for years. Right, the whole issue with the ends of the DNA. Exactly. Historically, the entire field has focused almost exclusively on the 5 or 5 prime ends of the DNA fragments.
2:29Because DNA has a directionality, right? Kind of like a one-way street. One end is the 5 foot and the other is the 3 foot end. Spot on. So if we're only looking at the 5 foot end, we're effectively ignoring half of the physical boundaries of every single molecule we sequence.
2:44Which sounds crazy. Why did the field settle for that for so long? Well, the reason is purely technological. I mean, it's not that the 3 foot end is biologically useless or anything. It's because of the conventional way we prepare double stranded DNA for sequencing.
2:57Because the sequencers are kind of picky, right? Extremely picky. When a tumor cell dies and its DNA spills into the bloodstream, it doesn't break cleanly. The fragmentation leaves these really ragged edges.
3:10Oh, so you have these single stranded overhangs where one strand of the double helix extends further than the other. Exactly. But sequencing machines historically anyway. They just don't handle jagged ends well.
3:21They need neat, blunt ends to attach the chemical adapters that basically allow the machine to read the genetic code. Okay, let's unpack this. So to get the DNA to lay flat and, you know, play nice with the sequencer, scientists use a chemical step called end repair.
3:36Right. The infamous end repair step. Yeah, where they artificially fill in the recessed 3 foot ends, or they use enzymes to just chop off the protrudi 3 foot ends until the whole molecule is blunt. That exactly what happens.
3:48So it's like investigating a crime scene and throwing away the scissors just because they don't fit perfectly in your standard evidence bag. That visual is so accurate. By forcing the DNA to fit our machines, we're actively erasing the biological footprint of whatever enzyme actually cut the DNA in the 1st place.
4:04Wow. We're literally tampering with the evidence. We really are. We've been doing what the researchers call one-dimensional fragmentomics. We looking at a highly sanitized, manipulated version of the evidence.
4:17And I guess for a long time, that was enough to find certain prominent tumor mutations, right? It was. But the team in Hong Kong recognized that we were hitting a wall with that approach. So they asked a really fundamental question.
4:30What was it? They said, well, what if we just don't end repair? What if we figure out a way to look at the native ends just as they exist in the plasma? Which sounds simple in theory, but chemically, that has to be a huge hurdle.
4:42Oh, it's massive. Because if the sequencing machines require these blunt ends to attach adapters. How do you sequence a jagged molecule without fixing it first? Right. So the team actually developed 2 distinct solutions for this, and tracing their trial and error is just fascinating.
4:58Okay, let's hear the 1st one. The 1st solution is what they call 2 end sequencing. And the big breakthrough here was basically abandoning the double stranded structure altogether during the preparation phase.
5:09Wait, they just got rid of the double helix? Yeah. Instead of trying to blunt the edges of a double helix. They used heat and chemicals to denature the self-red DNA. They essentially peeled the 2 strands apart into single individual strands.
5:23Ah I see. Because once it's just a single strand of DNA floating there, the whole concept of an uneven overhang just disappears. Exactly. It's just a single string of letters. You don't have to worry about how it aligns with the partner strand anymore.
5:35That is so smart. The chemistry bears that out perfectly too. Once they had those single strands, they could attach their sequencing adapters directly to the ends without any end repair process at all.
5:46So they preserved the native ends. Right. This preserved the absolute native 5 foot N, which they call EM5, and the absolutely native 3 foot N, which they call EM3, but then, you know, they went a step further computationally.
5:59Oh, I think I can guess where this is going. If you map that preserved single strand back to the human reference genome. You know exactly where the strand begins and ends, which means you also know what letters used to be there before the enzyme made the cut.
6:13It's like finding a torn page and using a pristine master copy of the book to figure out exactly what words were printed right next to the tear. That's a great way to put it. And what's fascinating here is just how much hidden information that reveals.
6:26They deduce the sequences immediately flanking the cut in the reference genome. Okay, so the surrounding neighborhood, basically. Right. The sequence upstream of the 5 foot cut became known as the pre-end motif, or pre, and the sequence downstream of the 3 foot cut became the post-end motif, or poem.
6:43So suddenly, instead of just looking at an artificially blunted 5 foot end, they have this rich constellation of 4 distinct data points for every single strand, Prim, EM5, EM3, and poem. Exactly, 4 separate pieces of data from one fragment.
7:00That is incredibly clever. But um, hold on. I'm thinking about the physics of peeling those strands apart. Okay, what's the catch? If you denature the DNA into single strands. Right at the beginning of your experiment.
7:11You destroy the relationship between the top strand and the bottom strand. You do. So you're reading them separately, but you don't actually know how they fit together in the original 3D space of that jagged cut.
7:22And you've hit on the exact limitation the researchers identified. The 2 end sequencing method provides excellent data on individual strands, but it breaks the physical molecular linkage of the native double spranded brake.
7:35Because you lose the shape. Exactly. To really understand how the enzymes are chopping up DNA in the blood. You need to see the exact shape of the overhanging cliff on both sides simultaneously. So they had to come up with another way.
7:48They did. Recognizing this limitation. The team pushed the boundaries again and introduced their ultimate innovation for end sequencing. Wait, I'm stuck here. If you can't end repair the DNA to make it blunt, and you can't peel it apart because you lose the 3D structure, how do you get a sequencing machine to read a fray double stranded mess?
8:08It seems impossible, right? Yeah, you'd need something that grabs onto the jagged ends perfectly, no matter what shape they are. Which is exactly why they had to invent a totally new tool. They design these structures called stem loop adapters.
8:19Stem loop adapters. Okay, what do those look like? Picture a hairpin made of DNA. In the open end of this hairpin, they engineered random single stranded overhangs of all different lengths and all different sequences.
8:31Oh wow. Yeah, they synthesize 1000000s of variations of these and mix them in a test tube with the native, jagged, cell-free DNA. So it's basically a massive molecular puzzle. These 1000000s of adapters are floating around acting like, I don't know, molecular Velcro.
8:48Molecular Velcro is exactly what it is. They just bounce around until they bump into a jagged self-free DNA fragment that perfectly matches their specific overhang. And the elegant part is that biology just does the sorting for them.
9:00The adapters naturally bind or hybridize to the complementary native ends of the tumor DNA. That's wild. So they just lock into place? Yep. Once they lock into place with a perfect match, the researchers introduce an enzyme that permanently glues them together.
9:15This process is called ligation, and it seals the structure into a continuous closed circular loop of DNA. Okay, but there's a vulnerability here, isn't there? What do you mean? Well, how do you guarantee this sequencer only reads the perfectly matched circles and not all the leftover garbage in the tube?
9:30There must be 1000000s of mismatched pieces floating around. You're totally right. You would need to clean it up. Don't they use exonucleses for that? Exonucleses are like little molecular Pac-Man, right?
9:41They eat DNA, but they can only start chewing if they find a loose open end. That's brilliant, yes. The mechanism relies on exactly that principle, because the perfectly matched molecules are now securely closed circles.
9:55They't have any loose ends. So the Pac-Man can't eat them. Exactly. The exonucleus enzymes chew up all the unlegated adapters, any damaged DNA, and any incomplete fragments. When the dust settles, you are left with only high fidelity completely intact circular molecules.
10:11And then you sequence those perfect circles. Yes. They sent those perfect circles to a pack biosingle molecule real time or SMRT sequencer. Oh, so specialized machine. Very specialized. This specific machine is designed to read around a circle multiple times, decoding the barcode on the adapter and the exact sequence of all 4 ends of the original double stranded molecules simultaneously.
10:34Wow. So you've gone through all this trouble to design stem loops, use exit nucleses to chew up the garbage and sequence these perfect circles. A lot of work. Yeah. Right. And if the cuts were just random damage from the physical stress of floating in the blood.
10:48This would all be a massive waste of time. Did they actually find a biological pattern? They found a massive biological footprint. Really? Oh, yeah. The 1st major discovery was that these precise points where the DNA is cut are absolutely not random, they are being purposefully sevied by specific enzymes in the body.
11:07Who could they prove that? They proved this by studying genetically modified mice. By comparing healthy wild type mice to mice that were genetically engineered to be missing certain nucleus enzymes? They could actually watch the fragmentation patterns change.
11:21Oh, that's incredibly definitive. It is. They proved that an enzyme called DNA C1L3 is the primary shearer generating these specific cuts. It leaves a highly recognizable signature on those prim and poem motifs.
11:35Okay, so knowing that a specific enzyme named DNAC1L3 is making the cuts, is great basic biology. But, um, how does this translate to helping a patient? Like if I'm worried about cancer, how does this footprint matter?
11:48This is where the clinical data comes into play. The team gathered blood plasma from patients with hepatocellular carcinoma. That's HECC, right? A type of liver cancer. Yes, HEC, which is a very common and aggressive type of liver cancer.
12:02They compare that plasma to healthy individuals, and importantly to patients with hepatitis B. Oh, because hepatitis B is a major risk factor for developing HCC. Exactly. So distinguishing between a patient who just has the virus and a patient who has actively developed a tumor is a massive diagnostic challenge.
12:17And if a liver tumor is growing, its cells are constantly dying, releasing DNA into the blood. So that tumor micro environment might alter how enzymes behave. Do the nucleus footprints actually look different in the cancer patients?
12:29The differences were stark. The end motifs in the cancer patient showed a highly distinct pattern compared to the healthy and hepatitis B groups. Wow. But, you know, the researchers didn't rely on just one metric.
12:42They combine the variables, they looked at prim, EM5, EM3, and poem, and they analyze this data across different sizes of DNA fragments. Different sizes. Yeah, they specifically focused on fragment populations around 52 nucleotides and 166 nucleotides in length.
13:00Oh, let's define that for a second. The 166 nucleotide length isn't just some arbitrary number. No, it's very specific. Right, because inside our cells, our DNA wraps around these protein complexes called nucleosomes, kind of like thread wrapping around a spool.
13:15A 166 nucleotide string is roughly the length of DNA wrapped around one of these coarse pools, plus a little bit of the linker DNA connecting it to the next school. And that biological consistency is exactly what you need to train a computer.
13:27Or they use machine learning. They did. When they took those predictable fragment sizes and fed this multidimensional motif data into a machine learning model, the diagnostic power was staggering. Their model achieved an area under the curve or AUC of .95 for detecting HCC.
13:45Okay, just to provide some context on that metric for everyone. An AUC of one.0 represents a perfect, flawless diagnostic test. And an AUC of .5 is basically a coin toss. So .95 is incredibly high for a non-invasive blood test.
14:00It means the model rarely misses a cancer case, and it rarely causes a false alarm. It is exceptional performance, but the true potential of this deep dive emerged when they kept pushing the data. But here's where it gets really interesting.
14:12They introduced a concept called 3 foot fragima. Yes. We previously knew about 5 foot fragment from this exact same Hong Kong team where fragmentation patterns at the 5 foot end correlate with DNA methylation.
14:23Which is super important for cancer. Right. To be clear, methylation refers to the epigenetic state of the DNA, meaning the tiny chemical tags sitting on top of the DNA sequence that turns specific genes on or off.
14:34Cancer massively alters these tags to fuel its own growth. Exactly. So the researchers naturally wondered, does the native 3 foot end, the part we used to throw away, also show us these methylation tags.
14:46And the foreign sequencing proved that it absolutely does. Really? Yeah, they discovered that the 3 foot ends strongly correlate with DNA methylation. Specifically, they found that the 3 foot ends most frequently terminate exactly one nucleotide before a methylated CPG site.
15:03Oh, wow. Let's break down a CPG site. That's a specific spot on the DNA, where a cytocene letter is next to a guanine letter, which is the exact place these chemical methylation tags like to attach. That's right So if the DNA usually breaks right before one of these tag sites.
15:18It's almost like the enzyme is scanning down the DNA strand, and when it hits that bulky chemical methylation tag, it acts like a physical speed bump. A speed bump is a perfect analogy. The enzyme gets blocked and makes the cut right there.
15:30Just one letter before the tag. And if we connect this to the bigger picture, it means the physical jagged shape of the DNA fragment floating in your blood is directly telling us the epigenetic state of the tumor it came from.
15:41We don't even need complex secondary chemical test to find the methylation. Exactly. The brake itself prints right to it. When they incorporated this new 3 foot fragima methylation data into their machine learning model, the detection of HCC jumped from an AUC of 0.90 using their older 5 corn method, all the way up to an astonishing .97.
16:04That is a massive leap in diagnostic power. Just from looking at the side of the DNA fragment, that standard sequencing protocols, literally chop off and throw in the trash. It's wild, isn't it? And, you know, the insights didn't stop at diagnostics either.
16:17The 4 end sequencing data revealed a beautifully coordinated biological dance of enzymes in the blood. What do you mean by a dance? Well, we often think of self-free DNA fragmentation as a single chaotic event, but the data showed it's actually an organized multi-step clearance process.
16:33Oh, so it's a teamwork exercise between different enzymes. Exactly. The foreign data showed that our primary enzyme, DNA is E1L3, tends to cut both strands of the DNA simultaneously, but mostly on medium-sized fragments in the 130 to 200 base pair range.
16:46Okay, so it has a specific target size. Right. However, another enzyme, DFFB, acts very early in the cell death process. It chops up the massive, long genomic DNA fragments that are over 600 base pairs.
17:00Wow, okay. And then a 3rd enzyme, DNA C one, acts much further downstream, sweeping up the remaining pieces and making cuts on the tiny fragments under 130 base pair. I love that visual. So DFFB is basically the lumberjack cutting down the big trees in the forest.
17:14DNA's one L3 is at the sawmill turning the logs into usable boards, and DNA's one is the wood chipper turning the leftover scraps into mulch. That is exactly what is happening. And because of the foreign sequencing, preserving the whole fragment.
17:27We can actually see the simultaneous cut marks of all 3 of these workers on the circulating DNA. That completely reframes our fundamental understanding of cell death and clearance in the human body. It really does.
17:37And when the researchers isolated just the foreign motifs, these comprehensive nucleus footprints, and use them alone to distinguish between the healthy patients and the HCC patients. The AUC hit .98. Okay, an AUC of .98 is almost unheard of in early cancer detection.
17:53It's so high that the skeptic in me is immediately looking for the catch. Fair enough We're talking about incredibly complex foreign sequencing here. Denaturing, stem loops, exonucleus, cleanup, pack bioreading.
18:07Like, can a local clinic actually draw my blood tomorrow and do this? This raises an important question, and you are totally right to be skeptical of the immediate clinical timeline. So it's not ready for tomorrow.
18:18The short answer is no, not tomorrow. This study is a profound proof of concept, but there are strict limitations we have to acknowledge. First, the clinical sample size for the 4 end sequencing portion of the study was quite small.
18:29They used 10 HCC patients for that specific method. Ah, okay, that makes sense. 10 patients is enough to prove the underlying chemistry works and that the biological signal exists, but you need 1000s of patients across multiple demographics to build a robust FDA approved diagnostic model.
18:46Exactly. The 2nd major hurdle is the hardware technology itself. To get that perfectly circular long raid data, the team utilized the packed bio sequel to platform. Right, which we talked about earlier.
18:57Yeah. It is an incredibly accurate piece of machinery, but it has relatively low sequencing throughput for this kind of broad application. We're talking about generating less than 1000000 genetic reads per sample.
19:09And that's not enough. For routine, high confidence liquid biopsies in a clinical setting, you generally need 10s or 100s of 1000000s of reads to gather reliable statistics on rare tumor DNA. So right now, the fore-end method is just too slow and too expensive to be a routine blood test.
19:28Oh. How do they bridge that gap between a brilliant lab experiment and a viable hospital tool? The necessary next steps involves scaling up the hardware and adapting the chemistry. They need to transition to newer, higher throughput, long read systems, like the pack by a review platform.
19:44Alternatively, the researchers note they are already looking into adapting the foreign technique for broader use on short red platforms like aluminum. Alumina. everywhere. Yeah, aluminum machines dominate the clinical testing landscape globally right now because they process massive amounts of data very cheaply.
20:00If the team can bring the fidelity of the fore-end chemistry to eliminate us throughput scale, that is the Holy Grail. So the translation of the clinic is basically a matter of engineering and scale, not a flaw in the fundamental biology.
20:13The biology is shouting a clear signal. We just need to build a bigger microphone to capture it efficiently. That captures the current state of the field perfectly. We've proven the signal exists, and it's richer than we ever imagined.
20:25Now the race is on to capture it at scale. So, what does this all mean? To synthesize everything we've unpacked today by developing innovative techniques to analyze the raw, unaltered ends of self-free DNA, researchers have unlocked a multidimensional view of how our DNA breaks down in the bloodstream.
20:42A totally new perspective. Exactly. This holistic fragmentatomics approach reveals the hidden teamwork of DNA cutting enzymes and drastically boosts the accuracy of liquid biopsies for detecting cancer.
20:53It really is a paradigm shift in how we look at genomic waste in the blood. For decades, the jagged edges were treated as garbage to be trimmed away. Now we know those edges are a detailed incident report of the tissue's health.
21:07What does this mean for the future of routine blood tests, and could these specific nucleus footprints one day allow us to map the precise origins of autoimmune diseases like systemic lupus, where the body's clearance of dead DNA goes awry?
21:21It's a thrilling frontier, and something we'll be watching very closely. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
21:34If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
21:48Thanks for listening, and join us next time as we explore more science, based by base.