This episode summarizes a PNAS study introducing CLSS, a contrastive two-tower protein language model that coembeds domain sequences, structures, and subsequences into a shared 32-dimensional latent space. Trained self-supervised on one million ECOD domains, CLSS aligns sequence and structure modalities, yields compact embeddings that recapitulate ECOD and CATH hierarchies, outperforms several state-of-the-art PLMs on ProteinShake classification tasks, and powers an interactive viewer for exploring protein space.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So, you know, usually when we talk about mapping a space, um, there's this expectation of absolute precision, right?
0:15Right, yeah, like everything just fits. Exactly. You pull up a digital map on your phone, you see the streets, the buildings, the topography, and it all just seamlessly layers together into one, I mean, one objective reality.
0:27Because the physical world provides a ground truth, a street exists at a specific coordinate. It intersects with a specific building. So all the data modalities, whether that's visual, spatial, textual, they all point to the exact same reality without much friction.
0:42But then you step into the world of evolutionary biology. Oh yeah, totally different story. Right. Specifically, um, trying to map out a 4 billion year history of the protein universe. And that seamless layering.
0:54It just shatters. We're looking at a landscape that is incredibly massive and deeply fragmented, mainly because we are working with 2 completely different sets of data. Right, the sequences in the structures.
1:04Yeah. On one hand, we have the protein sequences, which are like the linear text, you know, the amino acid letters. And on the other hand, we have the protein structures, the actual three-dimensional architectures.
1:14And historically, getting these 2 data sets to align has been just a monumental failure. We're essentially dealing with an informational schism here. I mean, thanks to the genomics revolution, we have 1000000000s of sequences.
1:28Yeah. Yeah. And thanks to Cryo EM and prediction models like, uh, alpha fold. We now have a massive influx of 3D structures too, yet they exist in entirely different mathematical and conceptual paradigms.
1:42It's wild to think about. Imagine trying to navigate a foreign city using only a list of street names, but you have no idea what the buildings actually look like, or vice versa. You have detailed photographs of every single building, but absolutely no idea what streets they are on.
1:56That is a great way to put it. What really happens when you can finally fuse the blueprint, the sequence with physical building the structure into one mastermap. That fusion is the central mission of this deep dive.
2:07And achieving that fusion. Well, it requires breaking down silos that have existed in computational biology for decades. For the longest time, sequence and structure have been treated as 2 completely different languages describing the same evolutionary universe.
2:22Right, with very little robust translation between them. But before we dig into the actual mechanics of how this translation was finally achieved, um, we need to spotlight the researchers who orchestrated this breakthrough.
2:34Absolutely. Today we celebrate the work of Guyanai, Giriel Axel, Liam M. Longo, near Ben Tal, and Rachel Culenny. This was an incredible collaborative effort between the University of Haifa, Tel Aviv University, the Institute of Science, Tokyo, and the Blue Marble Space Institute of Science.
2:52Yeah, their collective work here fundamentally advances our understanding of protein space. I mean, they haven't just iterated on an existing tool. They've provided an entirely new mathematical framework for how we visualize evolutionary biology.
3:04To really appreciate the magnitude of this framework, though. We have to look at the baseline. Historically, mapping the protein universe was this entirely manual human driven endeavor. We rely on these meticulously curated databases, right?
3:20Like e-cod, the evolutionary classification of protein domains and cath. The human labor poured into those databases is staggering. Seriously, for decades, structural biologists have just been manually aligning 3D coordinates on screens, trying to identify topological similarity.
3:39Just grouping proteins by hand. Exactly. Grouping them into rigid hierarchies like architectures, topologies, families. And they have to do this because algorithms like blast, which just look at sequence alignment, they completely fail when proteins have diverged over 1000000000s of years.
3:53Right, because the sequences change so much. Yeah. Two proteins might share less than 10% of their sequence identity, yet they fold into the exact same beta barrel. The human eye can spot that structural homology, but a basic sequence aligner is just going to see static.
4:07However, you know, relying on human curation introduces a massive bottleneck. It's incredibly slow. And because it relies on placing a domain into a strict hierarchical bucket. It's inherently rigid. We end up with a catalog of domain, sure, but we often miss the microscopic evolutionary bread crumbs.
4:25And those breadcrumbs are arguably the most fascinating part of evolution. We're talking about smaller subdomain sequences. Like how small? Like fragments of maybe 30 or 40 residues, these tiny pieces that have been repurposed by evolution to bridge 2 completely different folding structures.
4:42A strict hierarchical database struggles to capture that kind of fluidity because, well, it categorizes the final product, the whole domain, rather than the evolutionary raw materials. Right. It's also complicated by the fact that the relationship between the sequence string and the final 3D shape is far from straightforward. We have this many to many relationship.
5:00You mentioned sequences diverging while keeping the same fold, but the inverse is also true, right? We have metamorphic proteins. Wait, if the sequence is essentially the recipe, how can the exact same recipe bake 2 completely different cakes?
5:13It's a great question. When you look at the biophysics, the primary sequence dictates the thermodynamic energy landscape of the protein. Most proteins have a single deep energy funnel, meaning they reliably fold into one stable state.
5:27But metamorphic proteins. They have a landscape with 2 distinct, roughly equal energy minimum. Oh wow. So they can fall into either one. Exactly. The environment like changes in pH, binding partners or solvent conditions, it tilts the landscape.
5:43That causes the exact same string of amino acids to flip from, say, an alpha helical confirmation right into a beta sheet. So the environment acts as the final arbiter of the structure, completely breaking the old dogma, the sequence solely dictates a single static shape.
5:58This level of complexity is exactly why human curation just can't keep up, and why the field turned to AIs, specifically protein language models or PLMs. But earlier PLMs walked right into the same trap, didn't they?
6:11They kept sequence and structure and totally separate silos. Yeah, AI development and biology essentially mirrored the historical schism. You had transformer models trained exclusively on massive databases of one D sequences, just learning the grammar of amino acids.
6:25Then, on the flip side, you had graph neural networks or 3D convolutional neural networks trained entirely on spatial coordinates. Even when models attempted to ingest both modalities, they didn't force a mathematical reconciliation.
6:40They just kind of kept them apart internally. Exactly. The AI's internal representation kept the sequence data mathematically distant from the structural data. Which brings us to the architecture that finally forced them together.
6:51The research team introduced CLSS, which stands for a contrastive learning sequence structure. And the engine driving this fusion is contrastive learning. Right, so contrastive learning relies on a specific mathematical penalty, often using something called an info NCE loss function.
7:08How does that work in practice? Well, you take an AI with a dual architecture and you present it with a sequence and its corresponding true structure. The model is mathematically rewarded for pulling the representations, the embeddings of those 2 inputs, closer together in a high dimensional space.
7:24Okay, so matching the correct pairs. Exactly. And simultaneously, you sample 1000s of incorrect structures and force the model to push those negative pairs far apart. So the team implemented this using a 2 tower architecture.
7:38One tower is dedicated to reading the sequence, and the other is dedicated to reading the structure. But the configuration of these towers is where the strategy really shines, right? Oh, definitely. For the sequence tower, they used ESM2, which is a trainable language model with about 35000000 parameters, but for the structure tower, they used a massive 1400000000 parameter model called ESM3, And crucially, they froze ESM3.
8:03That was the brilliant part. Freezing that 1400000000 parameter model creates an immovable mathematical anchor. ESM 3 has already internalized the physical laws of protein folding by processing 1000000s of struction.
8:15Right. It already knows how things are supposed to look. Exactly. If both towers were allowed to update their weights during training, the latent space could just collapse or drift unpredictably. By keeping the structural tower frozen, they created a stable target.
8:28They are essentially forcing the smaller, more nimble sequence model to adapt its internal geometry to perfectly match the deep structural truths held by ESM 3. And they train this entire system on 1000000 e-cod domains in a completely self-supervised manner.
8:45So no cheat sheets, no human labels, just raw sequences and their structural coordinates, but they went a step further than just feeding the AI full domains, didn't they? Yeah, they created a specific variant called CLSS sub.
8:57With this, they train the model on random contiguous subsequences ranging from just 20 to 60 residues long. And that decision directly targets those evolutionary breadcrumbs we were discussing earlier, because a full protein domain is typically, what, 100 to 200 residues?
9:11Usually, yeah. So a 20 to 60 residue fragment represents a subdomain motif, perhaps a specific helix turn helix configuration, or a localized binding loop. Okay, so by forcing the AI to match these tiny fragments across sequence and structure, it learns the fundamental alphabet of structural evolution, not just the final paragraphs.
9:31Beautifully said. It learns how small local sequences dictate local architecture, and that is absolutely critical for identifying distant evolutionary relationships that might share only a single structural motif.
9:44So is this like a language app where it forces you to match the English word to the Spanish word over and over until the AI intuitively understands they mean the exact same thing? It operates on that exact principle, yeah.
9:55But the translation happens in a deeply compressed latent space. The model takes these incredibly complex high-dimensional inputs, a sequence of letters in a 3D point cloud of atoms, and projects them both into a shared mathematical vector.
10:09And the size of that vector is just astonishing. They shrank all this biological complexity down into a 32 dimensional mathematical space. I mean, in the context of large language models, which routinely operate in 1000s of dimensions.
10:2232 dimensions is a severe bottleneck. It is, but that severe bottleneck is a feature, not a bug. Really? How so? Well, when you force a model to represent complex data in only 32 dimensions, you strip away all the noise.
10:35The model simply doesn't have the mathematical capacity to memorize species specific amino acid variations or, you know, minor coordinate fluctuations. Oh, I see. It has to focus on the core identity. Exactly.
10:47It is forced to retain only the absolute core signal. The fundamental topological and evolutionary identity of the protein. And the effectiveness of that bottle egg becomes undeniable when you look at the benchmarks.
10:58The team ran CLSS through the protein shaped benchmark suite to see how well it could predict structural labels from the SOP database. you know, categories like families, superfamily, fold, and class. And the results were striking.
11:11Yeah. They pitted CLSS against several state-of-the-art models, including using ESM 3 alone, Pro T 5, and a highly toutered model called ProTrek, and CLSS significantly outperformed all of them. The performance metrics definitely validate the architecture, but I think the most compelling evidence actually comes from the dimensionality reduction maps, they used TSNE projections to visualize this 32 dimensional space in a 2D format.
11:35Right, to actually see it. Yeah. And when you plot the embeddings generator by CLSS, the dots representing the sequence inputs and the dots representing the structural inputs perfectly overlap. They map the exact same coordinates.
11:47So they successfully fuse the maps. But when you look at the TSNE projection for ProTrek, the story is entirely different. Project was literally designed to co-embed these things too. Why did ProTrek end up with sequences and structures segregated in different corners of the map while CLSS actually got them to hold hands.
12:07It really comes down to the strictness of the objective function. You see, ProTrek attempted a highly ambitious multimodal co-embedding. It was trying to align sequence, structure, and textual descriptions all simultaneously.
12:21Oh, so it bit off more than it could chew. Pretty much, without the highly focused, punitive contrastive loss that CLSS applied exclusively between sequence and structure, protracts, optimization landscape found an easier route.
12:33Mathematically, it was just simpler for project to cluster all the sequence representations in one region and all the structural ones in another, rather than truly intertwining them. So it cheated, basically.
12:43Yeah, kind of. CLSS allowed no such escape. It forced absolute structural and sequential parity. And because CLSS achieved that true fusion projecting the entire protein universe into one cohesive map, we can suddenly visualize massive biological trends.
12:58When the researchers overlaid functional data onto this TSNE map, the clustering revealed stark physical realities. For example, the map is literally bisected by cofactor utilization. Right. And co-factors are the non-protein chemical compounds, like metals, vitamins or organic molecules, like ATP or heme, that many enzymes actually require to catalyze reactions.
13:20When you map cofactor usage across the CLSS space, a clear boundary emerges. One hemisphere is densely populated by proteins that heavily rely on co factors, and these are almost exclusively folds that mix alpha heeluses and beta sheets.
13:35So we're talking about alpha beta and alpha plus beta topologies. And the other hemisphere is dominated by all alpha or all beta proteins, and they show minimal cofactor usage. What is the physical chemistry driving that massive universe level split?
13:47It's essentially a matter of structural topology creating opportunity. When a protein chain alternates between a beta strand and an alpha helix, the chain has to repeatedly loop back on itself to form the core of the structure.
14:00Okay, creating folds. Yes. And those topological crossovers naturally generate deep clefts, cavities, and flexible loops at the ends of the beta sheets. These cavities are the perfect physical pockets for trapping and binding a small molecule cofactor.
14:15Oh that makes perfect sense. Yeah, and conversely, tightly packed, all alpha bundles or rigid all beta barrels. They often lack these deep natural crevices, making them far less suited for cofactor binding.
14:27The AI independently mapped out the relationship between global full topology and chemical functionality. That is incredible. And it captured even more granular trends too, specifically regarding zinc binding.
14:39Zinc binding clustered tightly in very specific regions of the map, strongly correlating with architectures described as having few secondary structure elements. So disordered proteins. Exactly. These are proteins that contain highly irregular, loop heavy or intrinsically disordered regions.
14:54They just don't have the rigid stability of a massive beta barrel. Instead, they utilize a central zinc ion as a sort of structural pin. The amino acid side chains coordinate around the zinc, and the metal ion holds the floppy protein architecture together.
15:09And the AI grouped these proteins, based purely on the contrastive loss of their sequences and structures, completely unaware of the human definitions of disordered or zinc binding, which honestly perfectly transitions to the broader implications of this model, because the AI was trained entirely without labels.
15:27It had no concept of an e-cod architecture or a sea cath topology. Right. Yet when you measure the mathematical distances between different proteins in the CLSS latent space, those distances perfectly mirror the human curated hierarchies.
15:41That's amazing. It really is. Domains that human experts categorize as belonging to the same family are clustered tightly together. Superfamilies are slightly broader and folds are broader still. The AI basically organically deduce the evolutionary rules that human structural biologists has spent decades piecing together.
16:00If this AI learned the human rules without being taught them, what happens when the AI disagrees with the human experts? Does that mean the AI might have just uncovered a missing evolutionary link we've been blind to?
16:13I think that is the most exciting frontier for this technology, without a doubt. If CLSS places 2 protein domains adjacent to each other in the latent space, but eCod, or TEF, lists them in entirely different structural families, it signals crypticomology.
16:27So a hidden connection. Exactly. It suggests that these 2 proteins, despite looking radically different to human analytical tools, actually share a deep evolutionary ancestor or have undergone some complex fragment repurposing event.
16:39It gives researchers a highly targeted coordinate to search for missing links in the evolutionary tree. And to make that search possible, the team developed a CLSS web viewer, effectively democratizing access to this 32 dimensional space.
16:54If you're a protein engineer or a structural biologist listening to this, the work so implications are massive. You don't have to spend huge amounts of computing power, running a novel sequence through alpha fold just to figure out what family it might belong to.
17:08No, you can just embed the sequence into a 32 dimensional vector in milliseconds and instantly search the database for its nearest structural neighbors. The acceleration for drug discovery and synthetic biology is just profound.
17:21Finding structural analogs for target proteins or tracing the evolutionary trajectory of an enzyme can now be done with simple vector mathematics. But you know, we must map out the boundaries of this tool.
17:31CLSS is incredibly powerful, yes, but it comes with a specific structural limitation regarding how it was trained. Right. What is the catch? The primary limitation is its reliance on eCod boundaries. The AI was trained strictly on isolated protein domains, so it does not inherently know how to process a massive multidomain protein chain as a single continuous input.
17:51So, a researcher can't just feed the model a raw 2000 residue sequence and expect a perfect map. You have to computationally cleave your protein into its constituent domains 1st and then embed those individual domains into the CLSS space.
18:04Exactly. It requires a domain centric approach to biology. But even within that domain level resolution, the model's ability to seamlessly align the language of sequences with the architecture of structures.
18:16It's a watershed moment for computational biology. It really provides the unified blueprint we've been missing for decades. Okay, let's bring this all together. By combining the sequence text and the 3D architecture of proteins through contrastive learning, the CLSS model has successfully generated a unified 32 dimensional map of the protein universe.
18:36This AI breakthrough not only aligns perfectly with decades of human expertise, but also exposes hidden evolutionary patterns and functional splits across all known proteins. What does this mean for our ability to look backward and reconstruct the very 1st proteins that sparked life on Earth, or to look forward and design entirely new enzymes from scratch.
18:56This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe your podcast app and leave a 5 star rating.
19:09If you'd like to support our work, use the donation link in the description. Now stay with us for an original track, create especially for this episode, and inspired by the article you've just heard about.
19:18Thanks for listening and join us next time as we explore more science based by base.