A deep structural and evolutionary analysis of 11,269 enzyme structures across Saccharomycotina reveals how metabolic context sculpts protein architecture. The study integrates AlphaFold2 models, proteomics and metabolic models to map hierarchical constraints on enzyme evolution.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Hey everyone, it is great to be here for another deep dive.
0:11So I want to start today by having you imagine something. If you were tasked with building, say, a fleet of 10,000 vehicles, you would probably use the most cost-effective, readily available steel you could find, right?
0:26Oh yeah, absolutely. You'd want to keep the production costs as low as possible. Exactly. But if you only needed to build 2 highly specialized vehicles to perform a really critical, dangerous job, you might just spare no expense and build them at a solid titanium.
0:42Right, because the budget totally changes when the stakes are that high. Yeah. And it turns out, for the last 4000000 years, the microscopic fungi sitting right now in your kitchen have been using this exact same corporate accounting strategy to build the machinery of life.
0:58Which is just a wild concept to wrap your head around. It really is. Yeah. So what really happens when the chemical demands of survival force a microscopic protein to completely redesign itself, just to save cellular energy, could understanding this ancient 400 million-year-old structural blueprint change how we engineer new medicines or combat cancer or even develop biofuels.
1:22That is the big question. It forces us to look deep into the architecture of biology and recognize that evolution operates essentially as a, well, as a ruthless economic process. How could this change our fundamental understanding of life's machinery?
1:35Because the microscopic world is governed by a constant balancing act between maintaining vital chemical functions and, you know, actually paying the energetic cost to build the physical structures that perform them.
1:45Today, we celebrate the work of Oliver Lenke. Benjamin Murray Heineke, Marcus Ralzer, and an international team spanning Charity Universitets Medas in Berlin, the Francis Crick Institute, Talmers University of Technology, and more, who have advanced our understanding of how metabolism shapes enzyme structures over evolutionary time.
2:04Their work was published in the journal, Nature, on July 9, 2025, under the title, The Role of Metabolism, in shaping enzyme structures, over 400 million years. And the problem they are addressing here is just of immense scale.
2:18For decades, the scientific community has understood that enzymes, you know, the specialized proteins that act as catalysts for chemical reactions in our bodies are subject to evolution. Right. They're like the main targets for most pharmaceutical drugs, right?
2:31Exactly. They are the absolute focus of modern bioengineering. However, enzymes do not just evolve in an isolated vacuum. They evolve as part of a massive, heavily interconnected metabolic network. Which is incredibly complex.
2:43It is, and until this research, our understanding of the global biochemical constraints that physically shape the 3D structures of these enzymes over 1000000s of years has been, quite frankly, pretty fragmented.
2:56Okay, let's unpack this network idea. Think of a metabolic network, like a vast city transit system. I like that analogy. Yeah, we are trying to understand how the layout of the roads, the metabolic network itself, forces the physical evolution of the specific types of vehicles, the enzymes that navigate them.
3:15Right. And the scope of the study is vast. It examines the sacramicotini yeast subfylum that represents 4000000 years of evolutionary divergence. Which is a massive chunk of time. Yeah. This family tree spans everything from the common baker's yeast you might use to make bread, sacramices, syravisier, all the way to opportunistic human pathogens like Candy Dalbicans.
3:37And the core tension, the researchers identified here is the realization that cells must constantly negotiate a trade-off. They have to balance the absolute necessity of chemical efficiency with the biological cost of actually manufacturing these enzymes.
3:50Because synthesizing proteins takes energy, right? Exactly. It requires significant cellular energy ATP. It requires raw materials, an organism needs its metabolic transit system to run flawlessly, but it also needs to build the vehicles that operate within it as cheaply as possible.
4:06Without the engine spontaneously failing, obviously. Right, exactly. Wait, let me push back on the methodology here, though. They're looking at 400 million years of microscopic architectural history. You obviously can't dig up ancient yeast fossils and put them under a microscope to look at their protein.
4:23No, definitely not. So the team used alpha fold 2, an artificial intelligence system to predict these 3D protein structures. But if alpha fold is trained on modern protein structures, aren't we just looking at AI hallucinations of what ancient yeast look like?
4:38That is a very fair point. Like, how do they know these 11,269 structural models are actually historically meaningful? That is the critical hurdle with computational biology. And the team addressed it by grounding their AI predictions in robust evolutionary frameworks and actual experimental data.
4:57Okay, so it wasn't just entirely AI generated guesswork. No, not at all. They didn't just ask the AI to imagine ancient proteins. They took the actual modern genetic sequences of 26 highly diverse youth species, +one outgroup species, schizocechromyces pombe, just to root the evolutionary tree.
5:17These modern species represent distinct branches that diverged at different points over the last 400000000 years. By combining alpha fold 2's predictions with experimentally verified structures, they mapped these 11,000 diverse structures back to a central reference point.
5:34Which was the baker's East. Exactly. The well studied baker's yeast. So they are taking all these different distinct variations that exist today and overlapping them to see what changed and what stayed the same across the evolutionary tree.
5:46Yes, precisely. And they quantified this using 2 highly specific metrics. First, they calculated the mapping ratio. The mapping ratio. Okay. What does that tell us? It tells us what percentage of a particular enzyme's 3D structure perfectly aligns with the reference structure.
6:01Like, does the overall scaffolding hold the same physical shape and space? Okay, that makes sense. And second, they calculated the conservation ratio. This measures how identical the actual underlying sequence of amino acids is within those physically aligned regions.
6:16I want to dig into why the 3D aspect is so important here. Because if we have the oneD genetic code, the sequence of amino acids, why bother calculating the 3D mapping ratio at all? Wow. I mean, a one D sequence tells us the exact parts list.
6:31doesn't the code basically give us the whole evolutionary story. Not quite. What's fascinating here is that a parts list tells you what components are present. Sure, but it tells you absolutely nothing about where they are physically located on the assembled machine or how they interact.
6:45A oneD sequence might show that an amino acid mutated, but a 3D structure reveals the context of that mutation. Is the mutated amino acid buried deep within the stable load bearing core of the enzyme? Or is it sitting out on the flexible water exposed surface?
7:01And those 2 locations probably play by totally different rules. Completely different sets of rules. 3D structures reveal physical constraints like surface flexibility versus core rigidity that are entirely invisible when you just read a linear genetic sequence.
7:15Okay, so if the overarching environment reshapes the whole transit system, do all the individual roads within it adapt in the same way? Because when you start layering 4000000 years of data over these 3D models, the environment clearly dictates the machinery.
7:31Yes, it really does. The paper actually shows that an organism's diet directly forces structural changes down at the level of individual enzymes. Here's where it gets really interesting. What did these 11,000 AI generated models actually reveal?
7:47Well, the researchers examine species that ferment glucose like our bakers yeast, and compared them to species that rely strictly on aerobic respiration, which requires oxygen to process food. They discovered massive structural divergence in enzymes related to central carbon metabolism between these 2 groups.
8:05A prime example is the enzyme KGD 2P, which functions within the TCA cycle. The TCA cycle is basically the central engine of cellular energy production. Exactly. the core engine. So how exactly does a diet change the physical structure of KGD2P?
8:19I mean, it's the same enzyme doing the same basic job in the TCA cycle. Why does it need to look different? It comes down to the environmental stress and the sheer volume of chemical traffic. Respirating yeasts rely heavily on the TCA cycle for energy, meaning KGD 2P operates under high metabolic flux.
8:36Meaning it's working in overdrive. Right. And it faces significant oxidative stress from oxygen by products. Its structure adapts to become more rigid and stable to handle that constant, heavy load without degrading.
8:48And the fermenting yeast. In contrast, fermenting yeasts use the TCA cycle much less frequently. Because the enzyme isn't under the same constant, intense demand, the evolutionary pressure relaxes. Oh, wow. Yeah, the structure of KGG2P in fermenters is allowed to drift, accumulating mutations that make it more flexible or alter its secondary binding affinities.
9:10A slight loss and stability just doesn't threaten the survival of the fermenting organism. That makes perfect sense. The fermenter's engine doesn't need to be built like a tank if it's only driven on weekends.
9:20Exactly. But what happens when we zoom in from the organism's diet down to specific metabolic pathways? It seems like some chemical reactions are locked down, while others are just constantly mutating.
9:32The data reveals a strict hierarchy based on chemical dependency. Not all pathways evolve equally. Core oxidor duct daces. Those are enzymes that facilitate the transfer of electrons to generate energy are incredibly conserved.
9:47They show very little structural change across the entire 400000000 year timeline. And the same is true for metal binding enzymes. So if enzymes are vehicles, the oxyoductuses are massive, rigid commercial freight trucks locked to a very specific highway.
10:03You can't just casually change the wheels on a freight truck without causing a massive pile up. No, you definitely cannot. Moving electrons around a cell is a highly volatile, dangerous process. It requires absolute precision, so evolution refuses to alter the structure, but then you have hydro laces.
10:19These are more like nimble motorcycles weaving through traffic. That is an excellent way to conceptualize it. Hydrolases are enzymes that use water molecules to break chemical bonds. The paper notes that hydrolases, along with enzymes involved in processes like lipid metabolism or glycosylation.
10:34Which is the complex process of attaching sugar molecules to proteins, right? Yes, exactly. Those are highly divergent. They mutate frequently and exhibit significant structural variety across different yeast species.
10:45Why does evolution allow the hydrolase motorcycle to be so flexible and modular while the oxidor ductase freight truck is completely locked down. It is entirely about the reaction mechanism. Hydro laces often operate without the need for cofactors.
11:00Cofactors being that. Co-factors are external helper molecules like vitamins or specific metal ions that many enzymes absolutely require to trigger a reaction. Oxidoroductases rely heavily on these cofactors.
11:12Okay, so they need that extra piece to function. Right, meaning their physical binding sites must remain perfectly shaped to hold both the target molecule and the co-factor simultaneously. Hydrolysis, relying primarily on ubiquitous water molecules, just don't have to maintain that rigid multipart binding architecture.
11:30So they have more freedom. Yes. This chemical independence gives them the evolutionary flexibility to mutate, adapt to new types of lipids or change their surface shapes without breaking their fundamental ability to catalyze a reaction.
11:44Okay, so if you are wondering why a cell would care so deeply about the physical shape of a protein, we have to look at the raw economics of abundance and biological cost. The researchers found a direct relationship between how abundant an enzyme is in a cell and the specific raw materials used to construct it.
12:01Biology operates like a ruthless accounting firm. Building proteins requires ATP, the basic currency of cellular energy. The cell must synthesize amino acids, and the energetic cost of these building blocks varies wildly.
12:14Some are cheap, summer expensive. Exactly. Some amino acids are biologically cheap to manufacture, while others are incredibly expensive. To put that in perspective, an amino acid like glycine has a tiny side change as a single hydrogen atom.
12:27It is very cheap to make, but an amino acid like tryptophan or phenoline, has a massive complex carbon ring structure that takes dozens of ATP molecules to assemble. The researchers found that high abundance enzymes, the ones the cell requires 10s of 1000s of copies of evolve to aggressively utilize those cheap amino acids, like glycine and alanine.
12:49Because it's all about volume, right? Right. The cell optimizes for cost, where the multiplier effect is highest. If an organism is manufacturing a 1000000 copies of a specific protein, substituting an expensive phenoline with a cheap alanine saves a massive amount of ATP.
13:03And over time, that really adds up. Over 1000000s of years that energy savings translates to a distinct survival and reproductive advantage. Conversely, low abundance enzymes, those where the cell only needs a handful of copies, exhibit much greater structural diversity and freely incorporate expensive complex amino acids.
13:21Because the total copy number is low, the energetic penalty for using premium building materials is basically negligible. But wait, I have a major question about this rule, because the paper highlights a wild exception that breaks this accounting model.
13:35The biosynthesis pathway for thiamine, which is vitamin B1. Oh, yes, the Theer P and Theer 5 P enzyme. Right. They are very low abundance in the cell. According to the economic rules we just established, they should be highly diverse and full of expensive amino acids.
13:50But instead they are rigidly conserved. They hardly change at all. Why didn't evolution just design a normal reusable enzyme for this? Because of the extreme chemical demands of synthesizing thiamine. These enzymes, the 4P and the 5P function as what biologists refer to as suicide enzymes.
14:09Wait, suicide enzyme? Yes. During the chemical reaction to build a vitamin, the enzyme physically donates an essential part of its own structure. It literally destroys a piece of its own active site to force the reaction forward.
14:20Wait, really? They destroy themselves. They're essentially single used tools. They break themselves to make the vitamin. Exactly. And precisely due to that sacrificial mechanism, the total cost of synthesizing the enzyme is intrinsically tied to the cost of the vitamin itself.
14:35Thymine is a vital cofactor. If the cell cannot produce it, the cell dies. Wow, okay. High stakes. Because the enzyme must perfectly execute a highly complex self-destructive reaction, the cost of a structural mistake is catastrophic.
14:51If a mutation makes the enzyme even slightly less efficient at aligning the molecules before it sacrifices itself, the entire energetic investment is wasted, and the cell starves. So evolution rigidly locks down their structure, despite their low abundance.
15:06That's incredible. It proves that evolutionary conservation is driven by the total cost of failure, not just simple production volume. Exactly right. Okay, so we've established that the diet shapes the network.
15:16The chemical mechanism dictates the flexibility, and the production volume dictates the cost of materials. But where exactly is all this evolutionary accounting happening on the physical body, the enzyme.
15:26That's where the alpha fold models come in again. By mapping the evolutionary data onto the physical 3D models, the researchers observed a very distinct geographical hierarchy. Different zones of an enzyme operate under completely different evolutionary rules.
15:41Like the core versus the surface. Right. The core of the enzyme, the internal structural scaffolding is highly conserved. It is densely packed with alanine, which is a cheap amino acid that naturally forms very stable, rigid structures called alpha helises.
15:57The core must remain stable to prevent the entire machine from collapsing. And the outside of the machine. The parts exposed to the cellular fluid. The surface evolves rapidly and is highly divergent across species.
16:09The surface is the primary area where the cell performs its cost optimization. Evolution constantly swaps in cheap, flexible amino acids like glycine or serene along the exterior. So the surface acts as an evolutionary buffer, mutating to save energetic pennies without disrupting the internal mechanics.
16:26Exactly. But the actual business end of the machine, the binding site, where the enzyme grabs a molecule and forces a chemical reaction to happen, plays by an entirely different set of rules. The binding sites are heavily protected against mutation, right?
16:39They are. And most importantly, they completely ignore the rules of energetic cost optimization. Oh really? Yes, the cell will utilize whatever expensive, bulky, energy intensive amino acids are necessary to perfectly shape that chemical pocket, regardless of the ATP cost.
16:57The surface of the enzyme mutates to save energy, but the binding site operates with an unlimited budget to ensure survival. So what does this all mean for the big picture of science? This brings us to a massive technological implication.
17:10Because these agive binding sides are so deeply protected and conserved over 400 million years. They form distinct 3D shapes that the paper refers to as clusters. The consistency of these highly conserved structural clusters led to a profound realization by the research team.
17:25They realize they can be used as predictive beacons. Because the geometry of a binding site is evolutionarily locked down. You can train machine learning algorithms to scan uncharacterized mysterious proteins and search for these ancient, expensive, conserved clusters.
17:42So they can use this to find out what unknown proteins actually do. Yes. They demonstrated that you can accurately predict totally unknown small molecule binding sites, or even the exact locations where 2 proteins interact simply by identifying these evolutionary bedrock clusters.
17:58That is wild. If we have a newly discovered protein, and we have no idea what it does. We don't have to spend years doing trial and error chemistry. We just ask an AI to find the cluster that evolution refused to change, and we immediately know where the interaction happens.
18:12It is a phenomenal tool for accelerating the discovery of novel drug targets. However, we must calmly view these capabilities through the lens of the study's limitations. Right. There are always limitations.
18:23The data is fundamentally anchored in the Sacramicatina yeast subfilm. While 4000000 years is an immense evolutionary time frame, it is still restricted to fungal biology. So it might not map perfectly to human biology, for instance.
18:38Right. Furthermore, as you alluded to earlier with AI models, the analysis relies heavily on alpha fold 2. While highly accurate for rigid structures like alpha helises, AI models still struggle to perfectly predict random coil regions.
18:53Random coils being quick. The highly flexible, unstructured loops on the surface of proteins. The structural predictions in those floppy regions naturally contain a higher margin of error. Okay, but imagine you were a bioengineer trying to solve a modern industrial problem.
19:06If the surface is constantly mutating to save energy, but the core and the binding sites stay exactly the same, could future bioengineers basically swap the surface of an enzyme to make cheap custom-built proteins for medicine or industry without breaking the chemical engine?
19:22That is the exact frontier this research opens up. By understanding that the surface is evolutionarily tolerant to cost saving mutations. Bioengineers have a blueprint for optimizing industrial enzymes.
19:33So you can make them more efficient to mass produce. Exactly. If you possess an enzyme that degrades environmental plastics, but it is too energetically expensive for bacteria to mass produce, you could theoretically rewrite its surface genetic code using only the cheapest amino acids.
19:50That is incredible. You create a streamlined, budget-friendly exterior while keeping the complex plastic eating binding site perfectly intact. It transforms biological observation into actionable synthetic biology.
20:03We started by looking at how fungi evolved in our kitchens, and we ended up with a molecular blueprint for custom building the machines of the future. It really highlights that evolution isn't just random mutation.
20:14Biology is the ultimate architect. constantly balancing a strict energy budget against the absolute necessity of keeping the chemical engine running. To sum it all up, enzyme evolution is a masterclass in biological economics and physical architecture.
20:29Over 4000000 years, the immense pressure of the metabolic network has forced these microscopic machines to adapt, designing cheap, flexible surfaces to save crucial cellular energy, all while rigidly protecting the expensive, highly complex binding sites that sustain life.
20:44What does this mean for how we might engineer synthetic life forms to solve future global energy crises? It's a question that is going to drive the next decade of research, for sure. This episode was based on an open access article under the CCBY 4.0 license.
20:59You can find a direct link to the paper and the license in our episode description. If you enjoy this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
21:10Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science based by base.