This study builds the Acute Leukemia Methylome Atlas (ALMA) from 3,314 patient methylomes and presents three models—ALMA Subtype, AML Epigenomic Risk, and a 38‑CpG signature—that classify WHO2022 subtypes and predict 5‑year survival. The authors also demonstrate a nanopore-based specimen‑to‑result workflow for combined genome and epigenome profiling.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So, you know, if you fracture your arm, the diagnostic processes, well, it's essentially a solved problem in medicine.
0:13The x-ray shows this jagged white line and the doctor just points and says, there it is. It's binary, broken or not broken? Yeah, it's very clean. Exactly. It's clean. And honestly, if you're the patient sitting in that room, that kind of mechanical certainty is, you know, comforting.
0:27Oh, it is comforting. I mean, we are really drawn to things that are visible and perfectly categorized. But when you step into the world of oncology, specifically acute myoid leukemia or AML, that comforting diagnostic certainty just evaporates.
0:42Yeah, it gets murky fast. incredibly murky. You're suddenly looking at a landscape that is, well, chaotic. AML is a devastating disease. It's aggressive, moving rapidly through the bone marrow and blood, and it carries very high morbidity and mortality.
0:56But the real villain making it so deadly is disease heterogeneity. Right. Decades of research have shown us that leukemic stem cells are elusive and, you know, constantly evolving. Yes, exactly. Like you can have 2 patients with what looks like the exact same type of leukemia under a microscope.
1:12But their bodies will respond to the same chemotherapy regimen in completely different ways. And currently, the medical community relies heavily on reading the DNA sequence, looking for specific genetic lesions or mutations to classify the disease.
1:28Yeah, reading the raw text, basically. But what happens when the literal text of the DNA looks somewhat normal, or at least ambiguous, yet the disease acting out in the patient's body is fiercely aggressive?
1:39I mean, we hit a wall. Which is why we have to fundamentally rethink what we are looking at. What really happens when we stop just reading the genetic instructions and start looking at the chemical highlighters drawn all over them?
1:51Oh, wow. love that analogy. Right. We are moving beyond just the raw sequence of the genome and stepping into a radically different layer of biology. The epigenome. And mapping the sheer complexity of how leukemic cells use the epigenome to survive.
2:05I mean, that is a monumental task. Today we celebrate the work of Francisco Marchi. Jatinder K-Lamba, and their vast network of collaborating institutions who have advanced our understanding of epigenomic diagnosis and prognosis in acute myelode leukemia.
2:19Yeah, this is a massive international effort. It spans the University of Florida, MIT, Absala University, St. Jude Children's Research Hospital, and many others. It's a huge collaboration. It really is.
2:30And for listeners following along. This deep dive covers the open access article titled Epigenomic diagnosis and prognosis of acute myloid leukemia. It was published online on July 29, 2025 in a nature portfolio journal.
2:44And the DOI is 10.1038, 701-467-025-620054. Perfect. So if you're a clinician listening to this, you know the dread of seeing something like complex karyotype on a patient's chart. Oh, absolutely. It basically means the chromosomes are a mess, but it doesn't give you a precise target to hit.
3:03We've known for a while that DNA methylation, which is a primary mechanism of the epigenome, drives the proliferation of AML. But for those who might not spend their days thinking about molecular biology, let's unpack this.
3:15It's like trying to understand a complex novel by only looking at the letters, while completely ignoring the bold text, the italics, and the red warning stamps. I mean, we're missing the context. That is a perfect way to visualize it.
3:28What's fascinating here is that we already know epigenomic integrity is compromised in both young and old AML patients. I mean, the world health organization actually recommends molecular classification over mere cell shape or morphology.
3:43But hospitals haven't actually been doing this, have they? No, because they haven't had a feasible way to actually do this at the bedside. The technology to measure 100s of 1000s of microscopic methyl Asian roadblocks practically hasn't existed.
3:55Especially not in a time frame that actually helps a rapidly declining patient. Exactly. And that is the exact problem this research team set out to solve. They didn't just publish a paper saying, you know, the epigenome is important.
4:07They built a massive diagnostic engine called the Acute Leukemia Methylum Atlas, or ALMA. A-LMA. Okay. Yeah, LMA is a triumph of data harmonization. They gathered 3314 patient samples from 11 different clinical trials.
4:23Wow, that's a huge data set. It is. And for every single one of those samples. They looked at 331,556 specific CPG mevelation sites. Wait, hold on. I need to push back on that for a second. So they took 330,000 data points from 1000s of patients.
4:39Isn't that an overwhelming amount of data? I mean, how do you compress that without losing the most critical biological signals? That data overload is exactly why epigenomics has been stranded in the research lab and kept out of the clinic. The data is just too vast for human curation.
4:54So they use a machine learning algorithm called PacMap. PacMap. And what does that do? It's an unsupervised dimension reduction algorithm. It compresses those 330,000 data points into just 5 dimensions for classification.
5:05Okay, 5 dimensions. How does that actually work mathematically? Well, it calculates the relationships between all those data points. It ensures that if 2 patients have highly similar overall epigenomes in the 330,000 dimension space, they will remain tightly clustered together in the new 5 dimensional space.
5:24It basically preserves the topology. So you keep the similar patients clumped together. And then you have to figure out what those clumps actually mean. Precisely. Once they had those five-dimensional coordinates, they fed them into a supervised machine learning algorithm called light GBM.
5:41It detects patterns invisible to human curation and draws borders around those clusters, learning exactly where the boundaries are between different leukemia subtypes. So the software is only half the battle, right?
5:52The physical implementation of this, getting the data out of the patient's body to feed the software, that's where the real magic happens. Oh, absolutely. They developed a specimen to result protocol using long read nanopore sequencing.
6:03And the mechanics of this are just mind blowing. They take just 200 microliters of blood or bone marrow. Which is tiny. Yeah, that's less than a 10th of a teaspoon. And they use nanopore sequencing to natively read the 5 millisec methylation and the genome simultaneously.
6:18In one pass. Exactly. Because different DNA bases have different physical sizes, they disrupt the electrical current in the nanopore in unique ways. And a methylated cytocene is physically fatter than an unmethylated one, so it blocks more of the ionic current.
6:33Right, right. Creating a completely distinct electrical disruption. So without needing any harsh chemical treatments, the machine reads both the genetic alphabet and the epigenomic highlighters at the exact same time.
6:45And the most critical part, they do this in about 48 hours. Which is revolutionary. Yeah, when a patient is crashing from aggressive leukemia, you don't have 3 weeks to wait for a centralized lab. You need actionable intelligence immediately.
6:59So the software is trained, the sequencer is running. Let's look at what the LMA subtype model actually produced. The diagnostic power was remarkable. LMA subtype accurately predicted 27 different WHO 2022 subtypes.
7:12Here's where it gets really interesting for me. In their discovery cohort. There were 840 patient samples where the standard clinical tests basically shrugged. The clinical annotations were frustratingly ambiguous, labels like normal karyotype.
7:27Basically saying, we aren't sure. Right. And this algorithm was able to confidently place those 840 unknown patients into specific diagnostic categories based on their epigenomic signatures. That is a game changer for treatment.
7:41It really is. And there was an unexpected insight in how they train the model, actually, including data from non-AML samples, like acute lymphoblastic leukemia, mildest plastic syndromes, and even normal controls that actually made the machine learning classifier better at predicting specific AML subtypes.
7:57Wait, why would cluttering the data with other diseases make it better? Because machine learning fundamentally operates by defining boundaries. You have to teach it what a specific AML subtype is not. By feeding it data from completely different blood disorders and healthy controls, the model learned to draw razor sharp borders around the diverse AML subtypes.
8:15That makes total sense. So the algorithm finally gives a definitive name to the disease for those 840 unknown patients. But does putting a name to it actually change their odds of survival? I mean defining the disease as step one, predicting the future of the disease is step two.
8:32That brings us to their prognostic tool, the AML epigenomic risk model, and the predictive power they achieved here is sobering. It accurately predicted five-year overall survival. What did the numbers look like for the high risk patients?
8:46High risk patients in the discovery cohort, at a hazard ratio of 4.40 year for mortality. Let's break that down. A hazard ratio of 4.4 means that at any given moment during that five-year period, a patient in the high risk group was over 4 times more likely to pass away than someone in the low risk group.
9:02Yes, that is a massive divergence in survival trajectories. And what makes that number even more clinically powerful is that the epigenomic risk was an independent predictor. Meaning it works even when standard clinical markers are telling a different story.
9:15Exactly. It remains significant even when adjusting for standard clinical markers like age and minimal residual disease or MRD1. Wait, so even with a negative MRD one? Yeah. In the clinic, if a patient's MRD one comes back negative, it means the microscope can't find any visible leukemic cells, it looks like a massive success, but the AML epigenomic risk model could look at patients who had a negative MRD one and still identify which of them had high risk epigenomes hiding underneath primed for relapse.
9:43Wow. So even when the standard tests give the all clear, the Epigenome is waving a red flag. Did they validate this outside of the historical data sets like in a real hospital? They did. Testing on 20 actual hospital patients confirmed high concordance with standard of care genomic variants.
10:00They ran a 48 hour nanoport protocol from scratch, and the epigenomic signatures predicted by the model matched what the hospital's pathology lab found. So if you're the clinician deciding on the next phase of treatment, What does this mean practically?
10:12We are moving from a system of educated guesses to precise 48 hour molecular blueprints. It's a complete shift. Patients currently placed in ambiguous standard risk groups could be correctly identified as high risk.
10:26And that qualifies them for earlier bone marrow transplants or closer follow-up. Right. It moved the entire clinical approach from being reactive to being proactive. But let's look at the logistics. Whole genome nanopore sequencing requires computational power and expertise.
10:41If this is going to actually save lives globally, it can't just be locked inside massive research hospitals. The researchers understood that limitation perfectly. To democratize the science. They created a targeted 38 CPG panel.
10:55Oh, just 38 sites. Yeah. By measuring just 38 specific methylation sites out of the 330,000. They were able to achieve a highly similar predictive capacity for five-year survival. That is huge. A 38 site panel is something smaller clinics could realistically run without needing whole genome sequencing.
11:13Okay, let's pause the victory lab for a 2nd because we have to talk about the blind spots. An AI algorithm is fundamentally a reflection of the data it consumes. So what does this all mean when the algorithm encounters something incredibly rare?
11:25If it hasn't seen enough training data for a specific, uniquely rare karyotype, doesn't it run the risk of misclassifying the patient? This raises an important question about generalizability, and it's a limitation the authors are very transparent about.
11:39The model did struggle significantly with ultra rare subtypes, like Down syndrome associated AML, or rare genomic fusions like fus.erg. Because it didn't have enough examples of those specific fusions to draw a dedicated boundary for them.
11:55Exactly. The misclassifications precisely highlight the urgent need for larger, publicly available data sets representing rare patient populations. Biology is vast. And our training data must catch up to ensure rare variants aren't dangerously miscategorized.
12:10Which is why the team released the models in an open source python package called Alma Classifier, right? Yes. It's a living tool. As more hospitals sequence their rare cases and feed that data back into the open source model.
12:22The algorithm will dynamically learn and sharpen those boundaries. So let's bring it all together. For decades, we have been trying to read the book of leukemia by only looking at the letters. By leveraging machine learning and nanopore sequencing, the acute leukemia methylome atlas proves that the epigenome can rapidly and accurately classify AML subtypes and predict five-year patient survival.
12:44With striking independence from traditional clinical markers, yes. Most importantly, it takes complex molecular diagnosis out of the theoretical realm and puts it directly into a 48 hour clinical reality.
12:55It leaves you with a staggering thought. What does this mean for the future of personalized oncology? And how long will it be before your epigenome is as routinely checked as your blood pressure? That is the big question.
13:06This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
13:20If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
13:30Thanks for listening and join us next time as we explore more science base by base.