SciPhy is a BEAST2-integrated Bayesian framework that models sequential CRISPR‑based insertion edits to jointly infer time-scaled single-cell lineage trees, editing dynamics, and population growth. The authors validate SciPhy on simulations and apply it to HEK293T monoclonal expansion and murine gastruloid datasets, showing improved tree and branch-length inference relative to UPGMA and enabling phylodynamic estimates of growth.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Yeah, thanks for tuning in, everyone.
0:10We have a really fascinating discussion lined up. We really do. Today, we are launching into a deep dive on some incredible new genomics research. And to get us started, I want you to just imagine for a second, trying to track the precise ancestry of every single cell in a living organism.
0:27Right, which is basically the ultimate biological puzzle. Exactly. I mean, how does one single embryonic cell divide and then differentiate into a complex body with 1000000000s of cells? Finding out how that happened historically required, you know, massive amounts of guesswork.
0:43It really did. But what if each cell carried like a tiny flight data recorder, right there in its DNA? Ooh, like a log book. One that permanently recorded every division and every developmental milestone in real time?
0:55Yeah. So what really happens when we can read a cells exact family tree down to the minute it divided? If we had that, we'd be able to see the exact history of how a tissue or an organ or, well, an entire organism built itself from the ground up.
1:11We wouldn't be guessing anymore, we would be reading the actual historical record of life as it formed. And the crazy thing is, we are getting remarkably close to actually doing this in the lab. We really are.
1:22The technology is advancing so fast right now. So today, we are exploring a massive leak forward in how scientists actually decode those tiny cellular logbooks. Today we celebrate the work of the researchers at Eth Zurich, the University of Washington, and their collaborating institutions, who have advanced or understanding of single cell lineage tracing and developmental biology.
1:42Absolutely. The computational framework they've built here is honestly going to fundamentally change how researchers extract meaning from genomic data. So, to really appreciate why this new research is such a big deal, we 1st need to understand how scientists have historically tried to build these cellular family trees and where those old methods fall short.
2:02Right. Because for a while now, scientists have used CRISPR based lineage tracing. Okay, and how did the early versions of that work? Well, the early methods basically work by introducing random genetic edits, like little insertions or deletions, into a cell's DNA over time.
2:18Like physical scars on the genome. Exactly. Like scars. Then the cell divides and it passes that scarred DNA barcode down to its daughter cells. But the major problem with those early methods was that the edits were what we call unordered, right?
2:32Yes, unordered. They just happened randomly across the target sites on the DNA. So when you sequence the DNA later, you can see all the edits. But you don't know when they happened? Right. Figuring out the exact chronological order they happened in is incredibly difficult.
2:45And that severely limits how highly resolved your family tree can be. Okay, so enter the massive upgrade. We recently saw the invention of sequential prime editing systems, like the one called DNA typewriter.
2:58Oh, DNA typewriter. This is a huge breakthrough for the field. Because these new systems act much more like a literal tape recorder, right? Yeah, they use sequential genome medity. It employs a prime editor, which is a cast 9 case fused to a reverse transcript ace, along with some special guide RNAs.
3:16And what does that machinery actually do to the DNA? It lays down short genetic sequences, which we call inserts in a specific irreversible order at target sites. Wait, irreversible. Yeah. When one site gets edited, the system physically destroys the current recognition sequence and creates a new one downstream.
3:33So it moves along the DNA tape in one direction. Oh, wow. So we finally have a perfectly ordered numbered list of edits. Exactly. Edit one happened, then edit two, then edit three. The chronological sequence is physically locked in.
3:47But here is the major roadblock that this new paper addresses. The mathematical tools we've been using to reconstruct the lineage trees from that data. Well, they simply haven't caught up with the technology.
3:58Not at all. The standard computational tool people have been using to build these trees from the DNA tapes is a distance-based clustering method called UPGMA. UP. Right. Since we're unweighted pair group method with arithmetic mean.
4:10And all UPGMA does is look at the pairwise distances between cells. So it's just asking, how mathematically similar are the DNA tapes of cell A and cell B? Yes. And then it clusters them together based on that simple distance.
4:24Okay, let's untack this. So we've upgraded to a high definition sequential camera, but we're still viewing the footage on an old black and white TV. Exactly. If we connect this to the bigger picture, that analogy captures the flaw perfectly.
4:39UPGMA is missing the biological mechanics entirely. Because it's fundamentally just a distance matrix method. Right. It looks at the final genetic sequences and calculates a simple edit distance. Like, how many mutations apart are these 2 cells?
4:53It ignores the timeline completely? It's mathematically blind to the sequential nature of the edits. Yeah, and furthermore, it's blind to the physical, biological mechanics of how those edits are actually made inside the living cell.
5:05Meaning it doesn't know that certain edits might happen faster than others. Exactly. Or that some specific insertions are more biologically common. It treats the data like a simple math puzzle, stripping away all that rich byological reality.
5:19So you engineer this brilliant biological tape recorder in the lab, and then your software just flattens all the sequential data. Which is super frustrating. So the researchers realize they needed to build a completely new framework from the ground up.
5:32One that actually speaks the language of sequential editing. Right. And they call their solutions, siphe. That stands for sequential cast 9 insertion-based phylogenetics. I like that. Yeah, it's a Besian phylogenetic framework.
5:46And they implemented it in a really well-known software called Beast 2. Okay, let's break that down for everyone, because if you're picturing simple clustering, This is a massive leap in complexity. Well, it's a totally different universe from UPGMA.
5:58Right. Unlike UPGMA, which just measures simple distance. Saipe uses what is called a mechanistic model. Yes. It tries to mathematically model the actual physical process of CRISPR Prime editing inside the cell.
6:11And to do this, it actively calculates two main biological variables, right? Right, simultaneously. First, it calculates the editing rate, which is sometimes called the clock rate. So literally how fast the castine enzyme is making cuts and introducing edits into the DNA tape over time.
6:27Exactly. And second, it calculates the insertion probability. Which means what, exactly? Well, when a cut is made, what is the specific likelihood of a particular genetic insert being added? Say, the sequence CAT versus a completely different sequence, like GCG.
6:43Ah, so it's acknowledging that this editing isn't just a totally random, perfectly uniform process. Right. It's biological, so there are biases. And to handle this immense complexity. Saiphi uses a continuous time Markov chain and Filsenstein's pruning algorithm.
6:59Okay, Felsenstein's algorithm. How does that actually work in the context of these cells? Because you only have the finalized cells at the end of the experiment, right? Yeah, you start with the data you actually sequenced.
7:08The cells in your dish today? Those are the tips of the branches of the family tree. The algorithm then moves backward down the tree, step by step, toward the root ancestor. At every single fork, it uses the Markov chain to ask a question.
7:22Which is? It asks, given the editing rates and the biological biases we know exist, what is the mathematical probability that this specific ancestor gave rise to these 2 daughter cells? Oh, wow. So it calculates that conditional likelihood for every possible tree shape moving backward through time.
7:40So what does this all mean? I mean, translating that algorithm, it sounds a bit like reverse engineering a recipe by looking at the baked cake. That is a perfect way to put it. It integrates over all possible hidden histories to find the most mathematically probable sequence of biological events.
7:56And by doing that, it actually calibrates the branches of the family tree to absolute time. Yes, it doesn't just say cell A is related to cell B. It says they split exactly 3 days ago. Which is incredible, but having built this complex biological algorithm, the researchers had to prove it actually worked better than the old standard, right?
8:14They did. They tested it on computer simulations, then a lab culture, and finally, a complex 3D organism. Let's start with the Insilica simulations. So they generated simulated data sets where they already knew the ground truth.
8:27The real family tree and the exact editing rates. And how did Syphe do? When they ran Saiphe against UPGMA and another tool called Tide Tree, Saiphe significantly outperformed them both. Wow. It was vastly better at capturing the correct tree shape, the topology, and crucially, it accurately captured the time between cell division.
8:48Which we call the branch links. And that makes total sense because UPJA doesn't even know what time is in this context. Exactly. It's totally blind to time. So what happened when they tested it on real living cells?
8:57They looked at a 25 day cell culture experiment using a human cell line called HEK 293T. And this was a monoclonal culture, right? Meaning 1000000s of cells all started from one single progenitor cell.
9:10Yep. And all these cells had those DNA typewriter tape recorders in them. From this, Syphe found 3 really striking things that UPGMA had completely missed. Okay, let's look at the 1st finding. You mentioned insertion probabilities earlier.
9:24Did they actually find a bias in the edits? They did. Syphe proved definitively that the edits are not random. They found that a specific insert, the sequence CAT, happened about 16% of the time. Okay, 16% But an insert like GCG happened less than one% of the time.
9:41Wow, 16% versus less than one%. That is a massive biological bias. Why would the prime editor strongly prefer writing CAT over GCG? It actually comes down to the physical biochemistry of the reverse transcriptase enzyme used in prime editing.
9:57The structural chemistry of the DNA itself? Yeah. GC rich sequences form very tight bonds. They have 3 hydrogen bonds instead of two. So in the RNA guide, these sequences can form secondary structures, like tight hair pin loops.
10:11And I'm guessing the enzyme doesn't like those loops. Exactly. When the reverse transcript case encounters a tight hairpin loop, it physically stalls, which drastically reduces the efficiency of that specific edit.
10:21That is fascinating. So, Syphe accounts for this biochemical bias to avoid drawing false family ties. Right. If 2 cells both have a CAT edit, Syphe knows it might just be because CAT is biologically really easy to write, not necessarily because the cells are closely related.
10:38Which GPGMA would totally get wrong. Okay, the 2nd major finding was about time, because Saiphe estimates branch links in real time, they could estimate the doubling time of the cell population, right?
10:48Yes. Saipe calculated a doubling time of 28 to 33 hours. And what did UPGMA say? UPGMA estimated a highly unrealistic doubling time of 12.5 to 14 hours. Is a 12 hour doubling time even biologically possible for these specific cells?
11:04It is not. Real world, documented HEK 293T doubling times, sit comfortably in the 24 to 30 plus hour range. So UPGMA's legacy distance model basically forced a mathematically compressed tree. Yeah, resulting in a biologically impossible growth rate.
11:21Whereas Saiphe nailed the biological reality. Amazing. Okay, the 3rd finding from the cell culture experiment. This was about the DNA takes themselves, right? Right. The cells in this experiment actually had multiple DNA tapes inside them.
11:31Most tapes had 5 target sites for editing. But some teams were truncated. They only had 4 target sites. And Saiphe noticed a difference between them. It revealed that the tapes with only 4 target sites were actually edited faster than the tapes with 5 target sites.
11:46Wait, how does the physical length of the tape change the recording speed? Well, integrating a massive synthetic cassette with 5 target sites into a cell's genome might trigger local chromatin compaction.
11:58Or the bulky cast 9 enzyme complex might just physically struggle to navigate the larger integrated sequence. Just because of sheer hysteric hindrance. too big. Right. The foresight tape is just slightly more accessible.
12:10The point is, Saiphe was sensitive enough to detect this subtle real-world biological constraint. Simple distance clustering would just blur right over it. Here's where it gets really interesting. But wait, if the data is so good, Why did they find that the sampling proportion in the cell culture was up to 16 times lower than the scientists originally assumed?
12:30I saw that in the supplementary data. Oh, yeah. The physical lab counts, estimated they were sampling about one in every 1200 cells. But the ciphe model output claimed it was between one and 5000 to one in 20,000.
12:44So how does a computational model override a physical lab count? Is the math just drifting from reality here? Well, that actually comes down to the messy reality of the lab environment. When scientists estimate a population size in a dish, they're usually doing a bulk calculation.
13:00Just a rough head count. Yeah. But biological cultures are chaotic. Cells clump together, making them difficult to count accurately. And more importantly, there are massive bottlenecks during the single cell sequencing preparation.
13:13Like what? Entire lineages might not survive the harsh chemical dissociation process. Or the microfluidic cell sorter might simply miss them. Oh, so the cells are physically lost before they ever reach the sequencing machine?
13:25Exactly. Ciphe calculates the effective population size based on the coalescent events in the tree. If the branches imply there must be a much larger unsampled population to mathematically make sense of the genetic diversity.
13:38The algorithm reports that. It forces the researchers to confront the physical biases and sell loss in their own assays. The software is literally auditing the lab notes. But, you know, growing HEK 293 T cells in a flat two-dimensional plastic dish is one thing.
13:54They are immortalized cells. They don't represent the chaotic 3D spatial reality of an embryo. Right, which necessitates a move to a more complex model, and that was their final test. The marine gastroloids.
14:07Yeah, they tracked a single mouse embryonic skim cell is agreeing to a 3D structure called a murine gastroloid over an 11 day period. And a gasteroid is an artificial clump of cells that mimics the very early dynamic stages of embryonic development, right?
14:20Exactly. Cells aren't just dividing in a gastroloid. They are moving, they're changing shape, and responding to complex spatial signals. Okay, so to test real developmental biology, they apply to chemical treatment at day seven, a CHIR poll.
14:33A CHIR polls, yes. CHIR is a small molecule that effectively mimics the want signaling pathway. And in early development, that signals the cells to break symmetry, right? To start elongating and forming a head to tail axis.
14:47Right. So the big question was, can Syphe actually see that symmetry breaking event reflected in the DNA tapes? And can it? It absolutely can. What's fascinating here is just by reading the genetic barcodes.
14:58Sifi detected a massive shift in the population's behavior. Oh, massive. Well, before the CHAR pulse, the cells were dividing rapidly about one division per day. But right after the chemical pulse was applied.
15:10Psyve detected that the population growth rate dropped dramatically down to .5 divisions per day. Wow, 50% drop, because proliferation slows down as cellular resources shift to large scale cytoskeletal rearrangements.
15:24Exactly. And the transcription of differentiation genes. The cell stopped focusing on rapid division and shifted their energy toward moving and differentiating. And the slowdown in growth perfectly matches physical imaging data from previous studies, right?
15:37Studies that actually filmed gastroloids growing under microscopes. It's a perfect match. Syphe saw the exact same biological event, entirely in the dark, purely from reading the genetic barcode data. If we can track precise growth slowdowns, calculate division rates, and uncover hidden lab biases just from reading these DNA tapes.
15:58The applications for this must be absolutely massive. Only are. What this paper does is establish a completely new philodynamic framework for study developmental biology. Pyodynamic. I like the sound of that Yeah, for a long time, we were just asking the basic question, which cell is related to which?
16:15But now, with a framework like siphe, we can move far beyond that. We can ask when exactly did they divide, and how did their environment change their growth? Exactly. It's turning a static family tree diagram into a dynamic, time lapse movie of development.
16:29It's incredible, but of course, no computational tool is perfect. We need to look at the limitations and the necessary next steps for the field here. Yeah, we have to be realistic. Beesian models, especially ones running continuous time Markov chains on 1000s of complex sequences.
16:44Well, they have a reputation for being computationally demanding. Right. The computing power required must be huge. That is limitation number one. It is incredibly computationally heavy. To read the analysis on that single subset of 1000 cells from the cell culture experiment.
16:59It took about 15 days of continuous computing time on a high performance scientific cluster. 15 days for just 1000 cells. And humans have trillions of cells, so we're going to need much bigger computers.
17:12We need massive optimization. The team acknowledges they need to improve the software for speed, perhaps using new mathematical approximations that scale better. Makes sense. Are there any biological limitations?
17:22Yes, the other major limitation is what we call tape loss or dropout. Meaning the DNA recorder just goes quiet. Yeah, sometimes the cell recognizes the integrated DNA typewriter cassette as foreign or unnecessary.
17:34So the cell's regulatory mechanisms methylate the DNA, packing it into dense heterochromatin, silencing it completely. Oh, so the prime editor can no longer access it to make edits. Right. Or during single cell RNE sequencing, the transcript just isn't captured by the chemical primers.
17:51And if a cell loses its tape, you lose its entire history. Exactly. And right now, Ziphi doesn't have a perfect way to handle this missing data without potentially biasing the overall results. So they need to build in a dropout model to handle missing data without biasing everything.
18:08This raises an important question. As these CRISPR recorders get more advanced. Our computational models must respect the biological reality of the data, flaws, and all. We can't just pretend the missing data doesn't exist.
18:20Absolutely. We have to model the absence of information just as rigorously as the presence of it. We're building incredibly sensitive instruments, and we need the software, interpreting them to be just as nuanced.
18:32That's the ultimate goal. a fundamental shift from just drawing lines between dots to actually understanding the living, breathing timeline of how a tissue builds itself. To wrap this up, Syphe proves that by treating genetic lineage tracing, not as a simple distance puzzle, but as a biological time-based process, we can accurately reconstruct the dynamic history of living tissues.
18:53This allows us to extract vital developmental measurements directly from a cell's DNA. It's a huge step forward for the field. What does this mean for the future of understanding complex diseases like cancer, where tracing the exact lineage and changing growth rates of a single rogue cell could be the key to stopping it?
19:11I think it opens doors we couldn't even knock on before. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
19:24If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
19:38Thanks for listening, and join us next time as we explore more science, based by base.