This episode summarizes Hawkins et al.'s presentation of the Single-Cell Pediatric Cancer Atlas (ScPCA) Portal, a publicly available resource that provides uniformly processed sc/snRNA-seq data and standardized metadata for pediatric tumors. The Portal hosts summarized expression data for over 700 samples across 55 pediatric cancer types, downloadable as SingleCellExperiment or AnnData objects and accompanied by QC reports, automated and curated cell-type annotations, and CNV estimates. The team also introduces scpca-nf, an open-source Nextflow workflow using alevin-fry for efficient, reproducible processing and support for additional modalities such as CITE-seq, cell hashing, bulk RNA-seq, and spatial data. The resource aims to accelerate pediatric cancer research by reducing reprocessing time and enabling cross-sample analyses.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine for a 2nd um, that you are trying to assemble this massively complex jigsaw puzzle.
0:14Right, but it's not exactly a picture of a peaceful landscape or anything. No, no, it's actually the biological picture of a rare childhood disease. And now, imagine that every single scientist in the world who is trying to help you build this puzzle is using, like, differently shaped puzzle pieces.
0:31Oh, that sounds like a nightmare. It is. They're following entirely different rule books, and to make matters even worse, Some of the most critical pieces are just, well, they're completely missing. Yeah.
0:42And sadly, that is the frustrating, incredibly fragmented reality of pediatric cancer research today. So how could standardizing these genetic puzzle pieces change the speed at which we find cures for children?
0:53What really happens when you finally unify the world's fragmented pediatric cancer data? I mean, it fundamentally changes the entire landscape of discovery. When data is scattered and processed in a dozen different ways, You just can't easily compare notes.
1:07You're stuck. Exactly. You end up with these brilliant scientists working in total isolation. They're unable to see the broader biological patterns that really only emerge when you look at things at a massive scale.
1:20Today we celebrate the work of Alex's Lemonade Stand Foundation, and the childhood cancer data lab, who have advanced our understanding of single cell transcriptomics in pediatric tumors. We are taking a deep dive into an incredible effort to solve this exact data crisis.
1:36Yeah, it's called a single cell pediatric cancer atlas, and an atlas like this is, well, it is desperately needed in the pediatric field. Because the adult cancer field already has things like this, right?
1:46Oh, absolutely. If you look at adult cancers like breast cancer or melanoma, there are massive, beautifully harmonized databases. Projects like the human cell Atlas or the Human Tumor Atlas Network are for this comprehensive standardized data that literally any researcher can tap into.
2:03Right. But pediatric cancer is, thankfully, much rarer than adult cancer. Which is obviously a wonderful thing for children. Of course. But I can totally see how that creates a massive hurdle for the research side of things.
2:16If a disease is rare, The data is going to be incredibly sparse. Exactly. Because it's rare, a single hospital might only see, I don't know, a handful of cases of a specific pediatric tumor each year. So the data just gets siloed.
2:30Right. It just sits on local servers and individual labs across the world. Yeah. And critically, it totally lacks standardization. Every lab has its own favorite computational tools, its own way of organizing files.
2:42Yeah. A recent analysis actually pointed out that this lack of standardization is essentially the biggest barrier to reusing pediatric single cell data. Let's unpack this core concept first. Why are we specifically talking about single cell data?
2:55Because you know, we've been sequencing tumors for decades at this point. So why isn't standard old school bulk RNA sequencing enough for these researchers to find new drug targets? Well, to understand that, we really have to look at how tumors actually behave in the body.
3:11Tubors are incredibly heterogeneous. Meaning they aren't just one uniform blob of identical bad cells. Precisely. They are highly complex, almost like corrupted organs, in a way. You have this whole micro environment made up of different tumor cell clones that are actually competing with each other.
3:29Wow, competing. Yeah. Plus, you have immune cells infiltrating the tumor trying to fight it off. You have structural cells, you have blood vessels feeding the mass. a whole ecosystem. So if you use bulk RNA sequencing.
3:42You are taking all those 1000000s of diverse cells, grinding them up together and just, what, measuring the average genetic expression of the whole mess? That's exactly. You get an average. And averages in biology can be incredibly deceiving.
3:54I always think of bulk RNA sequencing, like making a smoothie. Oh, I like that analogy. Right. Like, if you take a massive basket of fruit, strawberries, bananas, blueberries, and then one incredibly sour toxic grape, and you blend it all together, the smoothie is just going to taste generally sweet, the strawberries mask the poison.
4:14You've completely hidden the individual rogue grape that is actually ruining the batch. That is the perfect way to visualize it. Averaging out all the cells, completely hides the individual rogue cells.
4:26And in oncology. Those rogue cells are often the most important ones in the entire tumor. Because they're the ones driving the cancer. Exactly. They're the rare clones driving the aggressive growth, or the ones that happen to have a mutation that resists chemotherapy.
4:40Or they're the specific cells that break off and cause the cancer to metastasize. So to find them, you cannot look at the smoothie. No, you have to taste every single piece of fruit one by one. We need single cell or even single nucleus, resolution to measure the gene expression in each individual cell.
4:57Which brings us back to the primary bottleneck. Researchers need the single cell data to find cures, but it's totally fragmented across the globe. So how does this new initiative actually solve that? By building a centralized, standardized translation engine, basically.
5:13The team created the SCPCA portal. That's the single cell pediatric cancer atlas. Okay. And backing that portal is a custom open source computational pipeline called Speaky Key and F. They essentially built a universal rule book.
5:30The pipeline takes the raw genetic readouts from all these different labs and processes them uniformly. But wait, processing the individual genetic readouts of 1000000s of separate cells, that has to require a ridiculous amount of computing power, right?
5:43You aren't just reading one genome, you're reading 1000000s of them simultaneously. How do they do that without just melting the servers? That's where the engineering gets really clever. They made a highly innovative choice for their primary quantification tool.
5:57Instead of using the standard software that most of the industry uses, Their pipeline utilizes a tool called 11 fry. Elephant fry. Yeah, it's incredibly fast and highly memory frugal. Memory frugal, meaning it doesn't require massive supercomputers to run.
6:11How does a program process 1000000s of genetic sequences using less memory? It comes down to how the software maps the genetic fragments. Older tools try to take every little piece of RNA sequence and perfectly align it to a massive reference genome, which just eats up a huge amount of computational memory.
6:29It's like trying to find where a specific sentence belongs in encyclopedia by reading every single page from start to finish. Exactly. But this newer tool, a Levin Fry. It uses a mathematical shortcut.
6:41It maps the sequences conceptually. It recognizes the unique signature of the gene, without having to read the whole encyclopedia. How that's smart. Right. It's highly efficient, and their benchmarking showed it returned completely comparable results while drastically reducing the computational burden.
6:57Okay, that makes the process accessible, but it doesn't necessarily make the data clean. I'm trying to picture the physical reality of this here. To sequence single cells, you have to capture them in these microscopic individual droplets of fluid in a machine, right?
7:12Yep, usually using something like 10 X genomics technology. Right. So if you're looking at 1000000s of droplets, some of them have to be empty. And what about cells that are just dead or dying? If a cell is falling apart, isn't its genetic material just degrading into static?
7:27You've hit on one of the most difficult quality control problems in single cell biology. If you include empty droplets or dying cells, your data gets incredibly noisy. Yeah, I imagine. Now, the empty droplets are relatively easy to spot computationally because they just have this ambient genetic soup floating in them.
7:46But dead or dying sales are much trickier. Because they still look like cells. Right. How do you distinguish a fragile dying cell from a cell that is naturally just quiet and not producing a lot of RNA?
7:58Because if the cell membrane is breaking down, the RNA inside is just leaking out into the droplet. Exactly. When a cell is damaged, its outer membrane ruptures in the cytoplasm, along with most of the RNA leaks away.
8:09But here is the biological trick the pipeline uses, mitochondria. The powerhouse of the cell. The very same. Mitochondria have their own distinct, incredibly tough double membranes. Oh, wow. Yeah. So even when the mainsell is dying and leaking its contents, the mitochondria tend to stay intact, they hold on to their specific mitochondrial RNA. The pipeline uses a statistical tool called my NQC to model the proportion of mitochondrial RNA against the total number of genes detected in each cell.
8:39That's brilliant. So if you look at a droplet and see that, say, 80% of the genetic material in there belongs to mitochondria, you know the rest of the cell has already leaked away. It's a dead cell. Precisely.
8:51A suspiciously high percentage of mitochondrial reads is the signature of a compromised cell. The pipeline flags it and boots it out of the data set. It basically takes out the biological trash, leaving researchers with only high quality viable cells.
9:05So we filtered the data, and the pipeline organizes it, grouping similar cells together mathematically using things like PCA and UMP, but we still have a massive problem. We have 1000000s of healthy cells grouped up, but they don't wear name tags.
9:19How do we figure out who is who? How does a computer actually know it's looking at a T cell versus, say, a neuron? That is the cell type annotation phase. And you're right. You can't just guess. To ensure accuracy, the pipeline deploys 3 distinct automated tools to manilize and label every single cell.
9:36Three of them. Yeah. One tool, called singler, compares the cell to known standardized reference data sets. Another, cell assign, looks for specific lists of marker genes genes we know are only turned on in specific cell types.
9:52Okay, and the third. The 3rd is similarity. It uses a massive foundation model, essentially an AI trained on 1000000s of cells to predict the identity based on deeply learned patterns. Wow, having 3 different AI and reference tools analyzing every single cell sounds incredibly robust.
10:08But what happens when the tools disagree? That does happen. Right. What if the 1st tool says the cell is a highly specific CD4 positive alpha beta T cell, but the AI foundation model just broadly says T cell, and the 3rd tool guesses something slightly different?
10:22This is where the pipeline does something truly elegant. It relies on an ontology aware consensus. Ontology aware. Yeah. In biology, we have a standardized family tree of cell types called the cell ontology.
10:34It maps how cells relate to one another. If the 3 tools give conflicting, highly specific answers, the pipeline mathematically walks backward up that family tree. Interesting. It looks for the latest common ancestor that all the tools can actually agree on.
10:50So if the tools can't agree on the exact hyperspecific subtype of the cell, the pipeline forces them to step back and say, okay, let's compromise. We all agree it's definitely some kind of lymphocyte. Exactly.
11:02It finds the most specific label that holds true across the different predictions. It's a fantastic way to harmonize annotations and ensure that the final label attached to that cell is trustworthy, rather than just blindly trusting a single algorithms guess.
11:16Okay, here's where it gets really interesting. So we have this incredibly clean, well labeled, standardized data, and the sheer scale of what they've processed through this pipeline is staggering. really is.
11:27We are talking about over 700 samples across 55 diverse pediatric cancer types, leukemia, sarcomas, brain, and central nervous system tumors, even rare, solid tumors, like neuroblastoma, and Wilms tumor.
11:41It is an unprecedented collection for pediatric oncology. But, you know, applying those AI annotation tools to these tumors revealed a fascinating fundamental blind spot in the field. Wait, a blind spot in the AI?
11:54Yes. Those automated tools rely on reference data to make their predictions, but the reference databases they learn from almost exclusively contain data from normal, healthy human tissue. Ooh. Right. They know what a healthy blood cell looks like.
12:09They know what a healthy neuron looks like. They do not have robust reference data for the highly mutated, bizarre, chaotic cells found in rare pediatric tumors. Oh I see. It's like installing a high-tech security camera and programming it with the faces of every authorized employee in the building.
12:23Yeah exactly. It's fantastic at recognizing the staff. But when a burglar breaks in, the camera has no idea who it is because the burglar's face isn't in the database. So it just flags the burglar as unknown.
12:33That is spot on. And when you look at the data sets in this atlas. You see massive populations of cells labeled completely unknown, because the AI simply doesn't recognize their corrupted genetic signatures.
12:44So what did they do? Well, the researchers were faced with a dilemma. How do you definitively identify which of those unknown cells are the actual malignant tumor cells driving the cancer? Well if you don't know the burglar's face, you look for the broken window.
12:58I love that. And in Genomics, the broken window is structural damage to the DNA itself. Normal, healthy cells have a very stable genome. Cancer cells however, are chaotic. Their chromosomes shatter, duplicate and stitch themselves back together incorrectly.
13:14The pipeline looks for this damage. Using a tool called Infer CNV, which estimates copy number variations. Copy number variations. That's when large chunks of a chromosome, meaning 1000000s of letters of DNA are either deleted entirely or duplicated multiple times.
13:30Exactly. But wait, this is an RNA pipeline. We are looking at the messenger molecules, not the DNA itself. How does reading the Messenger RNA tell you that the actual structural DNA is broken? It's a brilliant piece of computational inference.
13:42Think of a chromosome as a massive factory floor, and the genes are the individual assembly lines producing MRNA messages. Okay, following you. If you measure the MRNA output of a cell, and you suddenly notice that absolutely 0 messages are coming from the assembly lines on the short arm of chromosome 17.
14:01Well, you can mathematically infer that that entire section of the factory has been demolished. The DNA must be physically missing. Oh, that makes perfect sense. The absence of the messenger proves the destruction of the source.
14:13Exactly. And certain pediatric cancers have very well-known canonical structural breaks. In neuroblastoma, for instance, we frequently see the loss of a chunk of chromosome 1 Q, the massive duplication of 11 Q, and the loss of 17 P.
14:26So if a cell is sitting in the data set and the AI originally labeled it as unknown, but this structural tool looks at its RNA and says, hey, this cell is missing the entire short arm of chromosome 17 and has double the output from chromosome 11.
14:40Then it is highly, highly likely to be your malignant neuroblastoma cell. It leverages structural inference to overcome the limitations of the AI. That is so clever. And to fix this blind spot for the future, they launch the open SCPCA project.
14:54This allows researchers around the world to manually curate and label these tumor cells, feeding that human expertise back into the public database, so future AI models can finally learn what a pediatric tumor cell actually looks like.
15:07The ingenuity there is just incredible, but this atlas didn't just stop at single cell RNA sequencing, did it? For some of these rare tumors, they actually included multiple types of sequencing data. What's fascinating here is the multimodal discoveries.
15:21I think it's one of the most critical takeaways from this entire initiative. Really? Yeah. For several projects, specifically looking at brain tumors, wilms, tumors, and osteosarcoma, the researchers had single cell or single nucleus RNA sequencing data, and they had standard smoothie style bulk RNA sequencing data taken from the exact same tumor samples.
15:41Oh, so they had the smoothie and the individual pieces of fruit from the exact same basket. Yes. And because they had both, they could cross-reference them. They computationally simulated what the bulk data should look like by adding up all their individual single cells.
15:55Then, they compare that simulation to the real physical bulk data to see if anything was hiding in the bulk that the single cell process somehow missed. Wait, but why would something be missing from the single cell data?
16:08We already established that single cell is the higher resolution, more precise method. It is more precise in its measurements. Yes. But we have to remember the physical reality of the laboratory process.
16:18Okay. To get single cells, you have to literally take a solid piece of tumor tissue and dissociate it. You use enzymes and mechanical force to rip the cells apart from each other. That is a violent process.
16:29Oh I see. Certain cell types are incredibly fragile, and they simply get shredded during dissociation. Furthermore, in many of these frozen pediatric samples. Scientists can't even extract whole cells.
16:40They have to strip away the delicate outer cytoplasm entirely and just sequence the sturdy nucleus of the cell, what we call single nucleus sequencing. And if you strip away the cytoplasm, you lose all the RNA that was flooding around outside the nucleus.
16:55Precisely. When the researchers compared the modalities, they found something critical, the single nucleus sequencing was systematically failing to capture certain vital immune cells. Wow, really? Yeah, in Wilms, tumors, and brain tumors, they found that monocytes, a crucial type of white blood cell that acts as an immune century in the tumor microenvironment, they were vastly underrepresented in the single nucleus data compared to the bulk data.
17:18That is a massive biological red flag. If a researcher only ran single nucleus sequencing. They would look at their beautifully standardized data and think, oh, there aren't many monocytes in this tumor.
17:29The immune system isn't infiltrating. Exactly. They would assume the tumor was an immunological desert. And they would be completely wrong. It's a technical artifact of the sequencing method, not the actual biology of the child's tumor.
17:42Right. This discovery proves that having both data set side by side in the portal is absolutely vital. It ensures scientists don't miss the full picture, just because one specific technology has a physical blind spot.
17:54So we have this massive portal. Over 700 samples uniformly processed, incredibly clever workarounds for identifying cancer cells, and multimodal data to cover our technological blind spots. How does this actually translate to the day-to-day life of a scientist trying to cure pediatric cancer?
18:13It fundamentally changes how they spend their time. Before this portal, a computational biologist, wanting to study a rare pediatric sarcoma, would have to negotiate data access agreements across 5 different hospitals.
18:26Sounds exhausting. It is. Then they'd spend months just doing tedious data wrangling, downloading raw genetic fragments, writing code to harmonize them, struggling with server memory limits, trying to figure out which pipeline the original lab used.
18:39They are spending all their time just trying to get the puzzle pieces out of the box before they can even look at the picture. Right. The portal does all that heavy computational lifting for them. It packages the data into standard, ready to use digital formats, like single cell experiment objects for R or and data objects for Python.
18:56A researcher can download these preorganized files and instantly skip the coding nightmare. They jump straight into the biology. They can start their morning hunting for new therapeutic targets, or investigating why a specific subclone of the tumor is resisting treatment.
19:10It's like skipping the prep work and going straight to the scientific discovery, but I have to ask, are there limitations to how this data is served up? I know computational biologists can get into intense philosophical debates about how data sets from different hospitals should be merged.
19:25That is a very perceptive point. While the portal does merge data across multiple samples within a project so you can look at them together. The pipeline deliberately avoids performing batch correction on those merged files.
19:37Wait, why avoid it? Wouldn't you want to smooth out all the technical differences? The batch effects between a sample process to New York versus one processed in London? That is exactly the philosophical debate.
19:48Batch correction is essentially the tension between cleaning data to make it usable versus accidentally scrubbing away the very biological anomalies you were trying to find. I see. If the portal applied a blanket, aggressive batch correction algorithm to everything, it might smooth out the data so much that it accidentally erases a subtle, true biological signal that a researcher is desperately looking for.
20:12The appropriate way to correct for batch effects depends entirely on the specific scientific question being asked. So they leave that choice to the individual scientist. That makes a lot of sense. You give them a clean, perfectly standardized foundation, but you let them decide how to actually build the house.
20:29So what comes next for this initiative. Well, the field of genomics is moving incredibly fast. Currently, they are focused heavily on single cell RNA sequencing, which tells us what genes are actively being transcribed.
20:41But they note the critical need to expand to other modalities, like scatacse, which measures chromatin accessibility. Chromatin accessibility, meaning how tightly the physical DNA is pulled up inside the nucleus.
20:53Exactly. If DNA is tightly spoiled, the genes are hidden and turned off, if it's unwound, the genes are accessible and turned on. Understanding that physical structure gives you a deeper layer of the tumor's operational manual.
21:07And because the pipeline they built is inherently modular, it is designed to adapt and incorporate those new technologies as they mature. It really is a masterclass in open science and democratization.
21:18You no longer need to be at a massive, incredibly well-funded academic center with a dedicated bioinformatics department just to look at rare tumor data. A scientist anywhere in the world with a laptop and an internet action can download these files and start looking for a cure.
21:32It finally levels the playing field against a set of diseases that are notoriously difficult to study. If we connect this to the bigger picture, it gives everyone, globally, the exact same high quality puzzle pieces.
21:43To sum this all up for our listeners, the single cell pediatric cancer atlas, successfully unifies and uniformly processes single cell data for over 55 rare childhood cancers. By standardizing the computational pipeline and making analysis ready files freely available to the global community, it removes massive bottlenecks in pediatric oncology research.
22:06And as we look at the power of this incredible resource, I leave you with this to ponder. What does this mean for the future of personalized medicine when we can finally compare the cellular fingerprints of rare diseases across the globe instantly?
22:20This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
22:33If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
22:42Thanks for listening and join us next time as we explore more science base by base.