0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. We talk a lot about engineering life, but that usually means, you know, tweaking or optimizing things that nature already built.
0:15Right, working with existing blueprints. But what if? What if we could design brand new biological components? I mean, entire genes that have never existed before. Not by trying to guess their structure, but just by telling an AI what context they should show up in.
0:31I mean, that's really the holy grail of synthetic biology, isn't it? For decades, designing a functional protein from scratch has been just incredibly hard. You're either making tiny changes to a known sequence or you're using these massive complex structural models, which can be super slow and often struggle.
0:50Exactly. And the space of what's possible for a functional gene, it's astronomical. Natural evolution has only explored a tiny little corner of it. So, today what we're diving into is some research that's found a totally new way to navigate that space.
1:03They've built this generative genomic model. It's called evo that can create functional de novo genes. And when we say de novo, we really mean it. These are genes that have 0 significant sequence or structural similarity to anything we know of, this isn't just modifying things.
1:21This is genuine creation. It's guided by the language of the genome itself. Hey, so before we get into the nuts and bolts of how they pulled this off, we really want to recognize the innovative work we're discussing today.
1:33Absolutely. We're celebrating the team, including a DDT merchant, Samuel H. King, Eric McGuinn, and Brian L. High. They're affiliated with Stanford University and the Arc Institute. And their work, it just, it fundamentally changes how we can think about generative genomics.
1:48It's a robust new framework. It really is. They basically moved the field from, uh, treating the genome like a dictionary of parts you can look up. To treating it like a language, a language you can actually use to write entirely new biological sentences.
2:01Exactly. Yeah. To really get the breakthrough, though, you have to start with the main problem. How do you tell an AI what function is? You might say, I want a new enzyme, but how do you specify what you want that piece of DNA to do inside a living cell?
2:16Most past attempts have relied on things like sequence similarity, which just limits you to things that look like stuff we already know. Precisely. And this is where the paper makes just a genius move.
2:29They borrow a concept, but not from biology, from linguistics. Distributional semantics, right. I've heard this in the context of large language models for, you know, human language. That's the one. The idea is simple.
2:42The meaning of a word is defined by the company it keeps, the other words around it in a sentence. Okay, so how does that map to a genome? Well, it maps perfectly onto how genes are organized in precariotes?
2:53So bacteria and archaea. Functionally related genes often sit right next to each other in these clusters called operons. Oh, right. So you'll have a set of genes for, say, breaking down a sugar and they're all lined up in a row.
3:04Exactly. Because they all contribute to the same biological pathway. For decades, scientists view this for guilt by association. If you find an unknown gene next to known metabolism genes, it probably has to do with metabolism.
3:20But they're flipping it. They're not using it to figure out what something is. They're using it to create something new. That is the critical shift. They trained a model on massive amounts of prokaryotic genomes, so the AI learns these multi-gene relationships, it learns the grammar.
3:34And then it can do function guided design just by using the neighboring genes as a prompt. Okay, so let's get into the tech. The tool itself is called evo one. 5.5. It's a huge genomic language model. And crucially, it was trained on entire genomes at single nucleotide resolution, not just little snippets.
3:51This lets it see that big picture context, those opera instructures that can be 1000s of bases long. So the method, this semantic design, it sounds almost simple when you describe it. In principle, it is.
4:02You give the model a DNA trumped. Let's say the 2 genes upstream and the 2 genes downstream from where you want a new function, then you just ask it to auto-complete the missing gene in the middle. You provide the functional neighborhood, and Evo fills in the blank with a gene that makes sense there.
4:16Mm-hmm, that's it. So their 1st step was just to prove that Eva was actually using the context and not just, you know, memorizing the most common gene for that spot. Right. So they did an auto-complete test on a really well-known gene, RPOS, which is vital for bacterial stress response.
4:31And even when they gave the model, only 30% of the input sequence. Only 30%. It achieved an 85% amino acid recovery rate. Wow. Yeah, that's incredibly accurate. It is. But here's the really clever insight.
4:46The part that goes beyond just memorization. When they looked at the DNA sequences evo generated, they saw massive nucleotide diversity. Wait, okay, so the protein output, the amino acids were the same, but the underlying DNA code was different.
5:00Very different. It was using all sorts of silent mutations. And that's the tell. It proves Evo isn't just spitting back something from its training data. Ah, I get it. It's synthesizing completely new DNA sequences that still code for the right protein.
5:16It's learned the redundancy in the genetic code. Yeah, it's operating in that dark matter of sequence space. Finding solutions that work, but that evolution may never have stumbled upon. So, Evo can synthesize novel sequences that fit a context.
5:30The real test, then, is actual creation. Did it work? It worked incredibly well. They targeted 3 really challenging, highly diverse functions, starting with defense systems. Okay. They begin with toxic antitoxin systems, TAs.
5:43These are these little self-destructed dormancy switches in bacteria, and they are famous for evolving super fast and having very little sequence conservation. A tough target. Very tough target. For the protein protein systems, the T2TAs, ebo generated a functional toxin in 4 different antitoxins.
5:59And get this, those antitoxins only had about 21 to 27% sequence identity to any known proteins. 21%. That's deep in what they call the twilight zone, right? Where you basically can't predict function from sequence anymore.
6:12You can't. So the fact that these worked is solid proof of de novo design. That's incredible, but the function was even crazier, wasn't it? It was. Two of the generated antitoxins, Evo AT2 and Evo AT4 showed multitoxin neutralizing activity.
6:27Meaning they didn't just work against one toxin. They rescued bacteria from 3 different natural toxins. Really, mass, and yobi. That kind of broad compatibility is not something you see very often in nature.
6:39It's like the AI uncovered a more fundamental or a more modular way of solving the problem. And it didn't stop with proteins. They went after RNA systems too. Yep, the T3TA systems. It designed a functional RNA, antitoxin, and a toxic protein.
6:53And again, that protein EvoT1 had no strong sequence or even predicted structural similarity to any known toxin. That the, uh, the real mic drop moment for novelty seems to be the anti-Crispers. Oh, absolutely.
7:05Anti-Crispers, or acres, are what viruses use to shut down the bacterial immune system. They're hyper diverse, constantly popping up as brand new inventions. They're the perfect test case. And the success rate. A robust 17% experimental success rate in generating functional acres against Picatas 9, which, for a de Novo design task with no special fine tuning, is extremely high.
7:28Okay, so how novel were they, really? How did they prove it? They did this really smart analysis on the 2 most novel ones. Evo Acre one and Evo Acre 2. They showed 0 significant sequence or structural similarity to anything known.
7:41Right. Then they did what's called a residue coverage analysis. They basically ask, if we had to build these AI proteins out of little snippets of known natural proteins. How many different snippets would we need?
7:52Like trying to solve a puzzle. Exactly. And for Evoaker one and two, it took fragments from 28 to 31 different natural proteins to explain their composition. Wow, so it's not just a remix. It's something genuinely new.
8:03Genuinely new. It's a level of novelty on par with proteins designed by these incredibly complex, specialized computational pipelines. But Evo did it with just a simple contextual prompt. That is a massive difference in effort.
8:16Huge. And this versatility led them to scale up. They created Sin Genome, which is a public database of over 120 billion base pairs, of AI generated DNA. 120 billion. 100 billion. It's derived from prompts covering 9000 different functional terms.
8:34And it's high quality pick. It mimics the statistical properties of real procaryotic genomes. And it's already proving useful, right? They used it to confirm a link between a previously mysterious protein domain and cyterchrome C.
8:48They did. So it's not just a database, it's a discovery engine. So if we step back, what does this actually mean for synthetic biology? How does this change the game? Well, it's a completely new orthogonal approach.
8:59You're not starting with structure anymore. You don't need a mechanistic hypothesis. You don't even need to do task specific fine tuning for every new function you want. You're getting access to parts of that functional sequence space that evolution just hasn't touched.
9:12Or parts that older methods would have just thrown out because, you know, the predicted structure looked weird or low confidence. Exactly. Who cares what the structure looks like if the function works in context?
9:24And for researchers, having seigenome is like, I mean, it's a massive pregenerated library for gene discovery. It saves countless hours. You can go search it right now for a function you're interested in.
9:36Okay, but it can't be perfect. There have to be limitations here. The generation is auto-regressive, right? Predicting one base after another. That's a crucial point. It can sometimes fall into repetitive sequences or produce non-functional hallucinations.
9:51So you still have to test everything in the lab. You absolutely do. This doesn't replace the bench. But what it does is it dramatically improves the quality and the novelty of your starting candidates.
10:01Instead of screening a 1000000 random variants, you're screening a 100 highly plausible, totally novel ones. And the other big limitation is the reliance on context, isn't it? This works so well in bacteria because of those neat little operons.
10:16Correct. Applying this to, say, a human genome where genes are spread out and separated by vast non-coding regions. That's a much bigger challenge. It's going to require the next generation of genomic language models.
10:28But those limitations just kind of point the way forward, don't they? It feels like this is just the beginning. We're learning to write biology. I think that's the perfect way to put it. To sum it up, this model, Evo, uses the company a gene keeps, its genomic context, to successfully design, completely functional, de novo proteins, and RNAs.
10:46It's generating novel anti-Crispers, multifunctional antitoxins, all without needing any prior structural or evolutionary information. It's opening up these huge unexplored territories of functional sequence space.
10:57So if we can use the language of the genome to rapidly build custom biological parts that nature hasn't even thought of yet. What fundamental discoveries are now within our grasp, just waiting to be found in a database like syndrome.
11:10It's really exciting question. This episode was based on an open access article under the CCBY4 license. You can find a direct link to the paper and the lessons in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
11:25If you'd like to support our work, use the donation link in the description. Now stay with us for an original track, create especially for this episode, and inspired by the article you've just heard about.
11:34Thanks for listening and join us next time as we explore more science based by base.