ProtoCloud is a self-explaining deep generative model that embeds single cells around cell-type-specific prototypes to deliver accurate, uncertainty-aware cell type annotation and gene-level explanations from raw UMI counts.
0:00Welcome to Base My Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine you're peering into the absolute most fundamental blueprint of your own body.
0:12Right, right down to the microscopic level. Exactly. And right now, single cell genomics is currently mapping out 1000000s of individual cells in the human body to build these massive tissue atlases. And to sort through all of that, they're relying heavily on deep learning AI to sort and annotate them.
0:31Which, I mean, makes sense because processing that volume of data manually would take lifetimes. Totally. But here is the surprising fact. These deep learning bottles are incredibly accurate, but they function as total black boxes.
0:44Yeah, they just spit out an answer and hide all their work. Right. So if an AI spots a rare disease causing cell in your tissue, but cannot explain why or how it made that decision, Can we fully trust it?
0:55How could this change the way we verify life-saving AI predictions in biology? It's a huge dilemma. I mean, we finally have the computing power to handle 1000000s of cells, but we just lack the transparency to verify the actual biological truth behind those predictions.
1:11When it comes to human health, just trust the algorithm isn't exactly comforting. No, not at all. scientifically unacceptable, really. Well, the good news is that a new model has fundamentally solved this black box problem.
1:24Today, we celebrate the work of Cayung Guo and Jury Ding from the University of British Columbia, who have advanced our understanding of single cell cell type annotation. Such an elegant solution they came up with too.
1:36It really is. The paper is called proto cloud. A prototypical, self-explaining model for single cell analysis, and it was published in 2026 in cell genomics. It's a game changer for the field. Okay, let's untack this because constructing comprehensive cellular atlases requires assigning cell types to 100s of 1000s of newly sequenced cells.
1:56Deep learning handles this scale well, but fails at transparency. Right, and that's the core scientific hurdle here. We essentially need dual explainability. Dual explainability, so 2 layers of proof. Exactly.
2:09First, we need cell level explainability. Like, how does this specific cell compared to reference populations? Where does it fit in the larger map? In the 2nd layer? Gene level explainability. We need to know which specific genes actually drove the AI's decision to classify it that way.
2:28It's like having a brilliant medical assistant who gives you the perfect diagnosis every time, but absolutely refuses to show you their math or lab work. That is a perfect analogy. Like you want to trust them because their track record is great, but as the doctor, you need to see the proof.
2:44Yeah, you need to know if they looked at the blood work or the x-ray, right? The complication here is that tissues are not just uniform blocks of one thing. They consist of dominant cell types, the stuff that makes up the bulk of the organ and extremely rare cell types.
2:58Like, uh, how rare are we talking? I'm going to take basifils in the blood. They're super important for your immune system, especially with allergies. But they make up less than one% of your circulating blood cells.
3:08Wow, less than one percent. Yeah. So a standard black box AI might just absorb those wear base of fills into a larger, more common immune cell group simply because it doesn't have to justify its logic to anyone.
3:19It just takes the lazy route. But wait, isn't the raw data itself also a huge mess? Like even before the AI looks at it? Oh, absolutely. Single cell RNA sequencing data is completely plagued by what we call batch effects.
3:32Batch effects. Yeah, technical noise. So if a lab runs a batch of cells on a Tuesday and another batch from the exact same tissue on a Friday... The machine calibration or whatever is slightly different.
3:43Exactly. Room temperature, chemical reagents, all of it. It creates technical noise that totally confuses standard models. The AI will often group cells together just based on the fact they were both processed on a Tuesday.
3:54So it thinks Tuesday is a biological cell type that's hilarious, but also terrible for science. It's huge problem, yeah. So how does ProtoCloud fix this secretive medical assistant issue? Well, Proto Clouds architecture is built on a deep generative model.
4:09Specifically, it uses a variational auto encoder or VAE. Okay, a VAE. So imagine that, like a digital hourglass, right? That's a good way to visualize it. You pour all this complex data into the top. In this case, the input is raw UMI gene counts.
4:23UMI's unique molecular identifiers, right? Like tiny molecular barcodes attached to every piece of RNA so we get a true count. Exactly. So the VAE forces all those raw UMI counts. through a narrow bottleneck in the middle of the hourglass.
4:39And that strips away the irrelevant details. Yes. It forces the AI to compress the info and learn only the essential underlying traits. That bottleneck area is what we call the latent space. Which brings us to their 1st big innovation, right?
4:53Instead of standard classification, where the AI just spits out a text label, protocloud embeds these cells into this low-dimensional latent space, this sort of map and organizes it around prototypes. By default, it sets up 6 prototypes per cell type.
5:07So when a new cell enters the system, protocol doesn't just guess a label. It classifies the cell based entirely on this mathematical similarity to those specific prototypes. So that gives us our cell level explainability immediately.
5:18We can just look at the map and see, oh, it's standing closest to prototype number four. Exactly. But as we said, we need gene level explainability too. Right. How did it get to prototype number four? And that's innovation number two.
5:29Prototypical relevance propagation or PRP. Here, Pete, got it. What PRP does is back propagate that cell prototype similarity score backward through the neural network. So like walking backward through a maze.
5:43Yes. It traces the decision all the way back to the start in a single backward pass. And as it does that, it assigns a relevant score to every single gene. Oh, wow. So it literally highlights which genes acted like magnets pulling the cell to that spot on the map.
5:59Precisely. Okay. But I have a pushback question here. With all this biological data and technical noise. How does the model keep everything from just clumping together into a useless mess? Like, think of organizing a massive library.
6:11You want to separate the actual subjects of the books, the biology, from the different colored covers, from various publishers, the batch effect. What's fascinating here is how they design proto Cloud's 3rd innovation.
6:22A disentangled latent space. Disentangle. Yeah, they use a 2 stage curriculum, and they literally split the latent space, that map we talked about in half. Wait, they just cut the map down the middle. Conceptually, yes.
6:36The 1st half is strictly for encoding the biological cell identity. That's where the prototypes live. The 2nd half is completely dedicated to absorbing the technical noise, those batch effects. So it puts the book subjects on one side of the room, and the colored covers on the other.
6:50Exactly. And to keep the biological side from just collapsing into a single point. They use something called an atomic loss. Atomic loss. That sounds more like quantum physics than biology. It is inspired by physics.
7:01It uses attractive and repulsive forces. The attractive force pulls similar cells tightly to their prototypes, but the repulsive force pushes different prototypes away from each other. So they don't just melt into one giant clump.
7:14Right. It ensures the prototypes represent truly diverse biological states. That is wildly clever. And the results are just staggering. Across 11 different data sets, proto cloud matched or beat 8 state of the art methods.
7:29We're talking about beating major industry standards like Surat, scanning the VI, and cell typist. Yeah. And crucially, they did a stress test that blew my mind. The researchers deliberately shuffled and corrupted 20% of the training labels.
7:42They basically lied to the AI. Which happens in the real world constantly. Humans mislabel things all the time. Exactly. But even with 20% bad data, Protocloud maintained over 93% accuracy. Because it relies on the prototypes and the physical distance on the map, not just memorizing the text labels humans give it.
8:01Right. And let's talk about those highly relevant genes or HRGs that the backward pass identifies. Oh this is where PRP really shines. In one of the data sets, a lung data set protocol was asked to separate mature B cells from naive B cells.
8:15Which are super closely related, right? Very closely related. Hard to distinguish. But protoclod cleanly separated them, based entirely on specific gene expression. It pointed directly to genes like CD79B and LY9 as the drivers.
8:30It showed its math. And this leads to something incredible. Annotation correction. ProtoCloud assigns a certain or ambiguous confidence score. Yeah, and this is where it starts actively fixing human mistakes.
8:42Right. There's this massive peripheral blood data set called PBMC 30 K. Protocoloud flag cells originally labeled by humans as CD4 plus T cells and confidently reannotated them as cytotoxic T cells. A standard model doing that would just be assumed broken.
8:59But proto cloud brought the receipts. It showed that the HRG, the highly relevant gene NKG 7 had an expression level of 3.248 in these cells. Which perfectly matches cytotoxic T cells. Right. If there were CD4 plus T cells, the expression should have been around .932.
9:15It caught the human error and proved it with the exact gene expression data. That is just phenomenal for real world research, especially you look at time core studies. Like the one they did on retinal ganglion cells.
9:25The RG. Yeah. This was an optic nerve crush injury study. Yeah, severe trauma. They used iterative training to track changing cell states over 14 days. When cells get injured like that, they lose their standard markers.
9:40So they just look totally unrecognizable to normal models. Exactly. In the original human annotations, a massive portion of the damaged cells were just labeled as unassigned. But Protocol figured them out.
9:53Yes, it confidently reannotated 30 to 70% of those previously unassigned cells. Wait, really? That much? Yeah, 30 to 70%. For instance, it showed that the survival rate of C1 RGCs was actually 5.6%. And the original human estimate was, what, one.
10:092%? Right. The cells weren't dying off as fast as we thought. They were just changing their gene expression and hiding from standard models. Here's where it gets really interesting, though. They applied proto cloud to an esophageal mucosal cell Atlas, the EOE Atlas.
10:23Uh, yes. EOE is a chronic inflammatory disease. This is a great test of finding extremely rare cells hiding in complex tissues. In protocols, successfully identified, incredibly rare, disease driving cells like PRDM 16 plus in Gritic cells and ILC 2s.
10:40And it did this across independent test cohorts, which is the real kicker. Right, because those specific rare cells were completely missed in the original manual annotations of those test cohorts. Protocloud found them hiding in the noise.
10:54This raises an important question, though, about limitations. Because, as amazing as it is, protofloud currently relies entirely on having an annotated reference data set to train on in the 1st place. It needs that biological dictionary to start with.
11:08Exactly. You can't just give it an alien tissue sample and ask it to figure it out from scratch. Makes sense. But what are the next steps? Well future steps involve extending this architecture beyond just RNA sequencing.
11:18Modalities like scatexec, which looks at chromatin accessibility. Oh, like how tightly the DNA is spooled. Right, or even multimodal data, looking at RNA and DNA physical state at the same time. So what does this all mean?
11:30Let's summarize. Protocloud transforms single cell annotation from a black fox prediction into a fully transparent evidence-based framework using cell prototypes and gene relevance. By cleanly disentangling biological truth from technical noise, it accurately identifies incredibly rare cells and automatically corrects human annotation errors.
11:53It's giving us back the ability to trust the science in the age of AI. What does this mean for the 1000000s of black box AI predictions, sitting in existing medical databases, just waiting to be fact-checked and corrected by self-explaining models?
12:06I mean, it's going to keep us busy for a very long time. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
12:17If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode, and inspired by the article you've just heard about.
12:30Thanks for listening, and join us next time as we explore more science based by base.