This episode summarizes HANCOCK, a monocentric multimodal dataset of 763 head and neck cancer patients combining demographics, structured pathology and blood data, surgery reports, whole-slide images (WSIs) and tissue microarrays (TMAs). The paper demonstrates that multimodal machine learning and multiple instance learning with histopathology foundation models improve prediction of recurrence and survival and that the dataset is publicly available for research.
0:28Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Today we're embarking on a deep dyes into a truly groundbreaking scientific development. It's something that could really revolutionize personalized medicine, especially, you know, for tough diseases like cancer.
0:44If you've ever felt just overwhelmed by the sheer volume of scientific info out there. Think of this as your shortcut, a way to understand how combining lots of different types of patient data is unlocking completely new insights.
0:56Yeah, and what's fascinating here is how stubbornly complex diseases like cancer still are. I mean, how much we still need to learn, despite all our progress, even within what we call the same cancer type, the, uh, the diversity between individual patients, it just presents huge challenges, we really, really need better, more precise tools to handle all that complexity.
1:16So, okay, just imagine this for a second. What if we could actually pull everything together? All the puzzle pieces of a patient's medical history. I'm talking images, yeah, blood tests, even the detailed note surgeons take, and then use advanced AI to predict their unique cancer journey, like with incredible accuracy.
1:35What kind of aha moments could that unlock? You know, for doctors and patients, especially with really aggressive cancers. It's a really compelling idea and for a very good reason. The real world problem is, well, it's stark.
1:47Take head and neck cancer. It's pretty common, but the prognosis, often quite poor. We're talking five-year survival rates that can be as low as 25%, maybe up to 60% in some cases. Our current diagnostic tools or treatment pathways.
2:00They're valuable, don't get me wrong, but they often just struggle with that huge variability we see from one patient to the next. And that variability, that's exactly what makes getting precise treatment so incredibly difficult.
2:09Yeah, and when you think about the sheer scale of effort needed to even try and tackle a problem like that. It's just immense. So today we celebrate the work of the team at Friedrich Alexander Universe Attat, Erlong and Nuremberg.
2:23The Bavarian Cancer Research Center, BZKF, and the University Hospital Airline, who have advanced our understanding of precision oncology in head and neck cancer. Oh, absolutely. You really can't overstate that achievement.
2:34Collecting and, uh, carefully curating real-world data all from one center monocentric from 100s of patients. That's a massive, massive undertaking. This isn't some small, clean lab data set. This is messy, robust, real world information.
2:50It truly reflects the challenges you see in actual clinical practice. Creating this data set, HNCOC. It's really a critical step forward for the whole research community. It gives us a resource that's just been missing.
3:01Okay, let's break this down a bit more. What exactly is head and neck cancer? And why has personalizing treatment been? Well, such a tough nut to crack until now? Sure. So head and neck cancer is actually the 7th most common type of cancer globally.
3:15It usually starts in what are called squamous cells. They line areas like your mouth, your throat, your larynx. And unfortunately, it often spreads to nearby lymph nodes, and that, um, that tends to make the prognosis worse.
3:28Diagnosis usually involves something called a pendoscopy, basically looking around with a scope plus a biopsy. And then pathologists look closely at the tissue samples, treatment, often surgery, may be followed by radiotherapy or radio chemotherapy, and those choices are mostly dieted by, you know, how big the tumor is, what stage it's at, right, but the challenge is, despite huge efforts like the cancer genome Atlas TCGA, which gave us tons of genetic information, very few predictive biomarkers are actually used routinely day-to-day for head and neck cancer.
3:59There are a couple of examples, sure. Like, we know HPV, the human pepiloma virus is linked to better outcomes in certain throat cancers or pharyngeal carcinomas, and PDL one expression. That helps predict who might respond to certain immunotherapies.
4:11It's actually the only one routinely use for that right now, but that's that's really about it. Just a couple of knowns in this vast sea of unknowns. And this brings us right back to the core problem, this paper tackles.
4:21The huge lack of large multimodal, meaning multiple types of data and publicly available data sets, existing data sets. They often fall short. Maybe they have too few cases like, I saw one study with only 288 for radio mix.
4:33Radio mix. That pulling data from images, right? Exactly. Quantitative features from scans or another study, just 122 cases for proteomix, which is looking at proteins. And crucially, they often completely lack other important data types, like detailed blood work results or maybe the full surgeon's reports, this kind of fragmentation, this scarcity of rich, diverse data.
4:54It's really been a roadblock for developing true precision oncology. Wow, okay. That really paints a picture of the challenge. So knowing these massive gaps existed. How did these researchers actually go about building a data set to address them?
5:06And what clever techniques did they use to, you know, make sense of all this incredibly complex information? Right. So they developed this thing called H-A-N-C-O-C-K stands for head and neck cancel data set.
5:18It's comprehensive, comes from a single sender, so it's monocentric, and it's retrospective looking back at past data. They pulled together data from 763 head and neck cancer patients, collected all the way from 2005 to 2019, and they aggregated this huge array of diverse real world data points, all into one single powerful resource.
5:36Exactly. And the real strength of Ankok is in all those different data types, the multimodal components. Each one adds a critical piece of the puzzle. So 1st you've got demographics, things like age the median with 61 sex, about 80% male, and smoking status around 72% reformer or, and from that text, they manage to extract standardized codes, ICD, and OPS codes that classify the cancer type and the procedures done.
5:58That's hard work, getting structured data from free text. And really critically, the data set includes incredibly rich histologic images. We're talking whole slide images, WSIs, these massive digital scans of tissue slides, available for over 700 patients for the main tumor, nearly 400 for lymph nodes, all stand with H and E, hematoxylin, and Eosyn, the standard workhorse stain in pathology, plus, they included tissue micro arrays, TMAs.
6:24These are small cores of tissue taken from specific spots, the tumor center, the edge where it's invading. And these TMAs weren't just stained with H and E, but also with a whole panel of special stains, immunohistochemistry or IHC to light up different immune cells.
6:36VD3, CD 8, CD 56, CD 68, CD 163, and important markers like PDL1 and MHC1. a treasure trove of microscopic detail about what's happening right there in the tumor environment. Okay, wow. That is an incredible amount of diverse information for each patient.
6:53So, how did they then bring AI into the picture to make sense of it all? You mentioned their core approach was creating these multimodal patient vectors. Can you unpack that a bit? What does that actually mean and how does it help the AI?
7:07Sure thing. So essentially you take all those different data types who just talked about demographics, bloods, path reports, surgery notes, images, you encode each type, uh, basically translate it into a numerical format that the computer can understand.
7:21And then you concatenate them, essentially stack all these numerical representations together into one single long vector, like a big digital fingerprint for each patient. That's what they call an early fusion approach.
7:32You bring all the data together right at the start, before the main analysis. Now, these multimodal patient vectors, they're incredibly complex, super high dimensional. You can't just, you know, look at them and see patterns.
7:43Way too much information for our brains to handle directly. Exactly. So they used a clever technique called UM uniform manifold approximation and projection. It's a way to take that really high dimensional data and project it down into, say, 2 dimensions, something we can visualize on a graph.
8:00And what they found was really fascinating. Similar patients actually cluster together in this 2D map. For example, you could see groups forming based on known biology, like HPV positive throat cancer patients who also had lots of CD8 plus immune cells.
8:13They naturally grouped close together, it really showed that combining the data this way was capturing meaningful biological patterns. The AI was seeing something real. That is a neat way to visualize it.
8:25But okay, here's a classic AI challenge. Yeah. How did they make sure their models were actually robust, you know, generalizable, not just good at predicting on data that looked exactly like what they were trained on.
8:34Yeah, that's always the critical question, isn't it? And they tackled it with a really innovative method. They didn't just do one simple split of data into training and testing. They use something called a genetic algorithm, which is a type of optimization technique to specifically define 3 different train test data set splits.
8:53One was called in distribution, pretty similar patients in training and testing. Another was out of distribution. They deliberately picked the most dissimilar kind of outlier cases for the test set. And the 3rd split was intentionally biased.
9:07They put all the orpharyngeal carcinoma cases, that specific throat cancer type only into the test set, so the model had never seen them during training. Wow, so really stress testing the models. Exactly.
9:18This tough approach gives you a much more realistic idea of how the model will perform when it encounters truly new, unseen patients out there in the clinic. Okay, and what about analyzing those huge image files directly?
9:28The WSIs. That sounds incredibly computationally intensive. It is. These files are massive, but they used a technique called multiple instance learning, or MIL, specifically, they used a framework called clam.
9:41The beauty of MIL is that it lets the model figure out which parts of the huge image are actually important for the prediction without needing someone to manually draw boxes around every single cell or feature.
9:53Ah, so it learns to find the relevant bits itself. Precisely. It looks at the image as a collection or bag of smaller patches and learns which patches hold the key information. And they also pointed out that using pre-trained foundation models for histology, like one called UNI, gave really superior results for encoding the image features.
10:12These models have already learned a lot about histology images from massive devasets. Okay, so meticulous data, sophisticated AI. What did it all lead to? After all that work, what were the really big impactful results?
10:24What did Hancock actually reveal? Well, the core finding is pretty clear, that multimodal approach. It works. Combining all those diverse data types using machine learning, they used a random force classifier here was definitely superior for predicting key outcomes.
10:37Things like, is the cancer likely to recur? What's the patient's survival status? They reached a top average AUC score of 0.79 for both predictions. You see remind us quickly. Oh, right. Area under the curve.
10:49It's a common measure for how well a classification model performs. .5 is basically random chance. is perfect prediction. So .79. That's quite good, especially for complex clinical predictions, and definitely better than using single data types alone.
11:04And remember those tough test sets. They also confirmed the generalizability piece. Just as you'd expect when the models were tested on that unseen or a phrengeal cancer data set, the really out of distribution one, the prediction accuracy dropped.
11:18Now that's not a failure. It actually underscores how vital it is to test models rigorously. It shows you need representative data, and how performance might change when facing truly novel cases in the real world.
11:29Another really impactful finding came directly from the images, that multiple instance learning approach. It successfully predicted where the tumor originated, the localization just based on the whole slide image data alone.
11:39Like it could tell oral cavity samples apart, just by spotting gland tissue in the image patterns, and using those advanced image encoders like you and I. They got incredibly high AUCs for that, like .94 to .96.
11:50Wow, just from looking at the slide. Just from the slide data, yeah. It really shows the power hidden within those complex images if you have the right tools to analyze them. And they also found that integrating both the big whole slide images and the smaller tissue micro arrays, significantly boosted survival prediction.
12:06The combined image approach hit an average test AUC of 0.69 for survival. That was noticeably better than using just WSIs alone, which got .65 or just the TMAs, which only managed .52 on their own. Interestingly, they noted that the standard H and E stain tissues seem particularly predictive, which is good news because that's the most common stain used everywhere.
12:29And importantly, this wasn't just about finding new things. The data set also successfully reproduced stuff we already kind of knew. For instance, it confirmed the links between HPV status, PDO one expression, and patient outcomes.
12:41It also confirmed that having more CD3 and CD8 positive immune cells, those cancer fighting T cells in the tumor is linked to a lower chance of the cancer coming back. So it validates existing knowledge, too.
12:51Exactly, which is super important. It builds confidence that the data set is reliable and reflects known biology while also enabling new discoveries. Okay, let's unpack this. These aren't just abstract numbers.
13:02They represent tangible improvements in our ability to understand and potentially predict patient outcomes. That's incredibly impactful for real people facing these diagnoses. So what does this all mean for the future?
13:13How is hand cuco gonna push precision oncology forward? Well, the implications are pretty huge, honestly. First off, the data set itself is now publicly available. That's massive. It's set to become a really foundational resource for anyone doing multimodal machine learning research, not just in head and neck cancer, but really across precision oncology more broadly.
13:33It was explicitly built to help discover and validate new biomarkers, things that can guide more personalized treatments. The authors themselves also lay out a whole roadmap for future research based on this data set.
13:43Things like exploring that multiple instance learning for images even further, maybe integrating it with gene expression data transcript homex to get even richer information from image patches. They also suggest trying different ways to combine the data, maybe not just early fusion, but joint fusion, where the AI models are trained end to end on all data types simultaneously.
14:02And getting more out of those free text surgery reports using modern, large language models. That's another really exciting avenue. Plus, moving beyond simple yes no predictions. Can the models predict when an event might happen, or provide a more nuanced risk score?
14:15That requires different techniques, like regression. They also have data on other immune cell markers in Hansakuk already, CD 56, CD 163, MHC1 that haven't been fully analyzed yet, lots to explore there.
14:28And, of course, creating even more detailed image annotations. That would allow for even more targeted biomarker hunting within the images. Looking further out, they imagine integrating genomic or transcriptomic data directly into ancioscope to make it even more powerful long term, and maybe one of the most visionary possibilities this kind of data set opens up is the idea of digital twins for cancer patients.
14:48Digital twins. Like a complete virtual copy. sort of, yeah. A comprehensive digital model of a patient built from all this data, imagine being able to simulate different treatments on the digital twin 1st to see what might work best before you actually treat the real patient.
15:05That could dramatically improve clinical decisions. That's a truly incredible vision. But this raises an important question. Even with an amazing resource like HNCOV, there must be limitations, right? Challenges still to overcome.
15:17Yeah, nothing's ever perfect 1st go. Oh, absolutely. And authors are very upfront about that, which is good science. For instance, in this initial work. They mostly stuck to that early fusion approach for combining data.
15:28It worked well, but exploring other fusion strategies is definitely needed. Also, the current AI models are mostly doing binary classification yes no answers, moving towards predicting time to event, or continuous risk scores would be more clinically informative in many cases, and while the image annotations they had were valuable, they were described as sparse, rather than exhaustive, meaning not every single feature and every image was meticulously labeled, more detailed annotation would unlock even deeper insights.
15:57So yeah, plenty of work still to do to build on this fantastic foundation. Make sense. So if you had to boil down the absolute core message, the central insight from this deep dive into just 2 or 3 sentences, what would that be?
16:08I'd say the anti-Suck data set is really a landmark achievement in collecting and integrating complex real-world patient data for head and neck cancer. This rich multimodal resource, when combined with sophisticated AI shows enormous promise for improving how we predict patient outcomes, and, crucially, for speeding up the discovery of personalized biomarkers that can lead to much better cancer care.
16:32So, final thought then. What does this mean for how quickly we can transition to truly personalized and maybe even preventative cancer treatment strategies, not just for head and neck cancer, but potentially for complex diseases right across the board?
16:46This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
16:58If you'd like to support our work, use the donation link in the description. Thanks for listening and join us next time as we explore more science. base by base.