This episode reviews a study that evaluates how structure-based computational scores (AlphaMissense, FoldX DDG using PDB or AlphaFold2 templates, and RSA) compare with BayesDel for ACMG/AMP PP3/BP4 evidence in classifying BRCA1 missense variants. The authors used MAVE functional data and BRIDGES case-control validation to assess discrimination, evidence strength, and clinical risk association. Findings show AlphaMissense best discriminates functional impact and that combining AlphaMissense with DDG and RSA increases granularity of pathogenicity/benignity evidence. The study highlights that RSA strongly modulates benign evidence and that AlphaFold2 models can serve as DDG templates.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Glad to be here for another deep dive.
0:10Yeah, so I want you to imagine for a 2nd that you're sitting in a clinical geneticist's office. You know, you've just taken a genetic test to see if you have an inherited susceptibility to breast cancer.
0:22Right, specifically looking at the BRCA one gene, usually. Exactly. And you are sitting there expecting a very clear answer, like a binary, yes, you carry a risky mutation or no, your sequence looks standard, but instead of that, your doctor looks at the paperwork and just, uh, hands you a giant, agonizing question mark.
0:41Yeah, it's, I mean, it is a scenario that plays out in clinics around the world every single day. Which is terrifying for the patient. Completely. It represents this profound state of clinical limbo. Right, and they call it a variant of uncertain significance or a VUS.
0:55Basically, you have a mutation, so like a slight change in your DNA sequence, but the current state of science literally cannot tell you if that change is going to cause cancer, or if it's completely innocent.
1:06It's just a blank space in our medical knowledge. Right. And to understand why this happens, I always think it's helpful to, um, think of your DNA, like a massive generational cookbook. And a misense variant, which is what we're talking about today, is simply a single typo in one of those recipes.
1:23I love that analogy. So if the recipe says bake for 30 minutes, and the typo swaps the B for an M, so it says make for 30 minutes. The kick might taste a little different. Maybe the texture's off, but it still functions as a cake.
1:35Exactly. That's a harmless variant. But what if the typo changes the 3 to a nine? So now it says bake for 90 minutes. Oh, then the entire 3D structure just collapses in the oven. ruined. Right. It's completely destroyed.
1:48So how do doctors know which typo is the harmless make for 30 minutes and which one is the dangerous bake for 90 minutes? Well, that is the central challenge for the medical community right now. Knowing exactly which microscopic molecular typo causes the biological cake to collapse?
2:05And today, we're diving into a massive stack of recent clinical genomics research to figure out how, um, artificial intelligence and three-dimensional protein modeling are teaming up to finally eliminate this clinical limbo.
2:19It's honestly a fundamental shift. We're moving away from just reading the one-dimensional text of our DNA and starting to predict the actual 3D physical structure of the machines that the DNA built. Right, which brings us to the team behind this.
2:32Today we celebrate the work of Lubna Ramadan Moshadi, Miguel de La Jolla, and a really massive international collaborative team. Yeah, spanning multiple continents, actually. Yeah, institutions like the Hospital Clinico San Carlos and Spain, the University of Queensland to IMR Burko for in Australia, and Ambri genetics in the U.S.
2:50They've all advanced her understanding of classifying BRCA one genetic variants. And our discussion today is grounded entirely in their really incredible paper, which was published in the American Journal of Human Genetics on May 1st 2025.
3:04It's just so cool to see research that combines the expertise of, you know, molecular oncologists, biostatisticians, and clinical geneticists all working to solve this global bottleneck. It really takes a village to tackle a problem, this complex.
3:19Definitely. But before we get into their new AI solution. I think we need to understand how doctors are actually making these calls today. Like, what does the manual process look like when a clinician is trying to decide if a mutation is dangerous?
3:32Okay, so the current gold standard for classifying these genetic variants relies on a framework created by the American College of Medical Genetics and Genomics. Okay. Along with the association for molecular pathology.
3:45In the clinic, they just call it the ACM Jane P framing. Or CMGMP, got it. You can think of it almost like a courtroom trial for a specific genetic mutation. The clinician acts as the judge, and they're gathering different lines of evidence to build a case.
3:59So they're looking for points to tip the scales one way or the other. Exactly. They assign specific codes to the variant based on the evidence they find. So, for example, if computational tools predict that a mutation is dangerous.
4:10Like the prosecution's evidence. Right to the prosecution. They apply a code called PP3, which stands for pathogenic computational evidence. On the flip side, if the tools predict the mutation is harmless.
4:22Yes, the defense is evidence. Then they apply a code called BP4, which indicates benign computational evidence. And whichever side of the scale has the most evidence, dictates what the doctor actually tells the patient sitting in front of them.
4:36That's the goal. Yeah. Now, for the BRCA one gene specifically. The expert panel in charge of these guidelines currently relies heavily on a computational tool or all Baysdale. Baydell. Yeah. They use it to assign those PP3 and BP 4 codes.
4:49And Baysdell is used to look at mutations in 2 highly critical load bearing areas of the BRCA one protein, specifically the BRA and BRCT domains. But the thing is, based out primarily makes his predictions based on evolutionary conservation.
5:04Wait, evolutionary conservation, meaning like it looks at the sequence of DNA across different animals to see if it is changed over time. Yes, exactly. It adds to really symbol, but powerful question. If we look at a human, a mouse, a dog, and a fruit fly.
5:19Okay. Has this specific string of genetic code stayed exactly the same across 1000000s of years of evolution. Oh, wow. And if it hasn't changed at all across all those species, nature's basically screaming at us that this specific sequence is absolutely critical for survival.
5:35So any mutation there is likely catastrophic. Precisely. Okay, let's unpack this for a 2nd because evolutionary conservation makes perfect sense logically. But earlier, we talked about how proteins are physical 3D machines doing heavy lifting inside our cells.
5:50If the protein's physical shape is what actually matters. Why are our gold standard clinical guidelines basically treating these proteins like a flat one d string of text? Like, why has the actual 3D structure been sidelined in this whole courtroom trial?
6:05To be honest, it really just comes down to technological limitations. Historically, we simply didn't include 3D structure in these clinical frameworks because we didn't have high quality experimentally determined 3D structures for every protein.
6:19Because getting their structures is hard, right? Incredibly hard. Figuring out a protein shape in a laboratory requires a process called x-ray crystallography. It is painstakingly slow, it's incredibly expensive, and, well, some proteins just outright refuse to crystallize at all.
6:38They just won't cooperate. Right. So the clinical frameworks had to rely on what was abundant and cheap, which was one D sequence data. But relying only on the sequence means we miss the physical mechanics of why the mutation is breaking the protein.
6:51I mean, it's like trying to figure out why a car engine won't start just by reading the parts list without ever opening the hood to see how the parts actually connect. That is the perfect analogy. And the researchers behind this study recognize that relying solely on evolutionary sequence data leaves massive, massive gaps in our understanding.
7:08So they set out to prove that structural data, the actual physics of the protein, needs to be formally written into the clinical rule book. Okay, so they have this theory that 3D structure will give better answers.
7:19But to prove that, they need an answer key, right? They need a massive set of mutations where we already know, with absolute certainty, whether they cause cancer or not, so they can test their new structural models against it.
7:29Yes, and they utilized an incredible data set known as M-A-V-E, which stands for multiplexed assays of variant effect. It is honestly a monumental scientific achievement. What exactly is in it? The research team took 1638 specific misense variants in the BRCA1 gene that had already been physically tested in a wet lab.
7:51Wait, really? Someone physically tested over 1600 variants. Yes. Scientists had manually observed whether each of those mutated proteins function correctly, or if they failed and lost their function. This MAVE data set served as their absolute ground truth, their answer key.
8:07Okay, so they have this answer key of 1638 known mutations. The next logical step is to throw our newest computational models at those variants to see if the computers can grade the test correctly. Exactly.
8:20So what kind of tools did they use to assess the 3D structure? They utilize 3 distinct structural approaches. First, they used alphemous cents. I've heard of that. That's the Google Deep Mind AI, right?
8:32Yes, a highly advanced AI system built upon the architecture of alpha fold. Alphamism is trained to predict how likely a mutation is to cause disease by looking at both evolutionary patterns and the structural context of the protein.
8:46Okay, that's tool number one and the 2nd approach. The 2nd approach focused purely on thermodynamics. They used a tool called Foldex, specifically Foldex 5.0 to calculate the folding stability of the protein.
8:56Somo dynamics, so like heat and energy. Basically, they measured the physical strain on the molecule to see if a mutation would require too much energy for the protein to hold its proper shape. And they ran these thermodynamic calculations on both real laboratory structures and AI generated models from Alpha Fool 2.
9:13Okay, wait, before we get to the 3rd tool, I want to clarify something about these 1st two. Sure. If we are using AlphaSense, which is an AI, and we are also using AlphaFold AI models to run the thermodynamic stability calculations.
9:24Aren't we essentially just testing AI against AI here? Like, how do we know one isn't just copying the other? That is a very valid question, and it's a crucial distinction to make. We are actually using 2 fundamentally different types of engines here.
9:38Alphemous sense is a deep learning neural network. It has ingested 1000000s of protein sequences and structures, and it recognizes incredibly complex invisible patterns to output a likelihood score. But it doesn't actually know physics.
9:53It's pattern recognition on a massive scale. Correct. Full X, on the other hand, is a strict physics engine. It knows nothing about evolution or cancer or deep learning patterns. It literally just runs mathematical equations for thermodynamic physical strain.
10:10It's calculating the atomic forces pushing and pulling on the molecule. Oh I see. So the fact that FoldX is running its physics math on a 3D blueprint generated by alpha fold, doesn't make FoldX and AI tool.
10:22One predicts patterns while the other calculates raw physical stress. They are asking entirely different questions. of the mutation. That makes perfect sense. Okay, so we have pattern recognition and we have physics calculations.
10:34What was the 3rd tool? The 3rd approach was a metric called RSA, or relative solvent accessibility. This calculates the physical geography of the mutation. Geography, like where it lives on the protein.
10:46Exactly. Is the specific mutated amino acid buried deep inside the protective core of the protein, or is it exposed on the outside surface, interacting with the watery environment of the cell? Okay, this sounds a bit like a filled pastry to me.
10:59Yeah, like the buried amino acids are the jam filling, locked safely inside the dough, and the exposed amino acids are the crust on the outside, interacting with everything else in the bakery box. That is actually a very apt way to visualize it.
11:12yes So the researchers took these three concepts. Alphemous sense patterns, thermodynamic physical strain, and the crust versus filling location, and random against the MAVE answer key. The ultimate question being, did this new 3D structural approach actually perform better than the old sequence-based tool, Bezdel, that clinicians are currently using?
11:33The results were definitive. The structural approach significantly outperformed the sequence only tool. Really? By how much? Well, alphemisms proved incredibly adept at discriminating between harmless and broken variants.
11:45It achieved an area under the curve an ORC score of 0.93. Wow, that's really high. It is, but the most impactful metric for the patient sitting in the clinical geneticist office is how these tools affect that agonizing limbo state, the VUS category.
12:00Right. How many patients were they able to actually move out of the limbo category? When the researchers optimized the scoring thresholds using their new structural guidelines, only 5% of the variants fell into the uninformative, uncertain score range.
12:14Only 5%. Yeah. Compare that to the current standard with Baydell, where 14% of the variants were stuck in limbo. Wait, that is a massive reduction in uncertainty. That is essentially 9% more patients getting a definitive yes or no instead of a giant question mark.
12:29It's a huge step forward for clinical clarity. But I mean, science is rarely a perfect sweep, right? Did the models fail anywhere? Because earlier you mentioned the pastry analogy, the crust versus the filling, and I suspect the location of the mutation threw a wrench into the system somewhere.
12:45You're spot on. It did. And it actually led to one of the absolute most critical bombshell revelations in the entire paper. Oh, wow. What happened? The researchers discovered that these computational tools.
12:58Even the highly advanced alphemous ins and the standard Baysdell completely failed to provide reliable, benign evidence for mutations that sit on the surface of the protein. The crest in your analogy. Wait, wait, wait.
13:09So a protein can be evaluated by our most powerful AI. The AI can confidently say, hey, this mutation is completely harmless, the patient is fine, but if that mutation is on the outside crust of the protein, we actually cannot trust the computer's answer.
13:23We cannot trust it for benign calls, no. And we have to look at the biological mechanics to understand why. Okay, let's unpack that. Why does it f? Because these algorithms and physics engines are incredibly good at detecting, if a mutation will destabilize the proteins fold, like if it will melt the core of the molecule.
13:42If the jam filling leaks out. Right. But a mutation on the surface of the protein rarely ruins the core fold. The protein usually remains perfectly stable. So the computer sees a stable, well folded protein and just assumes it's functional, but stable doesn't necessarily mean it's doing its job.
13:59Proteins do not operate in a vacuum. BRCA one is a tumor suppressor protein. To repair damaged DNA and prevent cancer. It has to physically bind with other helper proteins. It has to perform a structural handshake.
14:13Ah, I see where this is going. Yeah, a mutation on the surface might leave the BRCA one protein completely stable, but it alters the surface chemistry just enough to ruin that critical handshake with its partner.
14:24So if it can't make the handshake, the DNA goes unrepaired, and the patient is at a high risk for cancer. And the computer completely missed it because it was only checking to see if the protein was structurally intact.
14:35Exactly. This is why human oversight and understanding the limitations of our tools is paramount. We absolutely must know the physical location of a mutation, the RSA, before we blindly accept a benign score from an algorithm.
14:50That is wild. That completely changes how we should interpret AI predictions in medicine. It's not just about the score. It's about knowing what the AI is actually capable of seeing. It all about context.
15:00Now, everything we've talked about so far is based on laboratory tests and computer simulations. Did the researchers validate these findings in the real world? Like, does this hold up when we look at actual human patients?
15:12They did. Which is what elevates this paper from a theoretical exercise to a clinically vital study. They cross referenced their structural findings against the bridge's data set. What's the bridges data set?
15:25It's a massive registry containing genetic data from over 50,000 population-based breast cancer cases. Oh, wow. So they check to see if the mutations the computer flagged as physically unstable actually translated to real women getting cancer.
15:41Yes. And they prove that the specific variants flagged is dangerous by these structural physics tools correspond directly to clinically actionable breast cancer risk levels. Incredible. The data showed odds ratios over four.
15:54That means patients carrying these specific, structurally unstable variants are more than 4 times as likely to develop breast cancer in their lifetime. It confirms that the physical strain calculated by the computer mirrors the biological reality in the human body.
16:09That is such powerful validation. Now, I want to circle back to something we touched on earlier regarding the methodology. Because this part fascinated me Sure, what's that? When they were running the thermodynamic physical strain calculations, you said they didn't just use x-ray crystallography structures from the lab, they also use structures generated by alpha fold.
16:28Yes, the AI models. How did the AI models perform compared to the painstakingly slow laboratory models? This represents a major paradigm shift for the field. The researchers found that the thermodynamic calculations ran just as accurately on the AI generated LFO 2 models as they did on the real experimental crystal structures derived from a laboratory.
16:49You're kidding. That is, I mean, it's like finding out that a hyper realistic video game physics engine is just as reliable for crash testing cars as actually building a real car, putting a dummy inside and crashing it into a concrete wall in a lab.
17:03That's a great way to put it. The implications for scale are just staggering. It saves an unbelievable amount of time and resources. Absolutely. It means that we can theoretically assess the structural stability of misense variants in proteins that we haven't even mapped in a physical laboratory yet.
17:19We don't have to weigh years for a crystal structure. If alpha fold can accurately model the protein's 3D architecture. We can immediately begin running physics simulations on it to look for vulnerabilities.
17:30But I'm guessing there's a limitation here, right? Because AI isn't a magic wand that solves biology overnight. Did the alpha fold models work perfectly out of the box for every single part of the BRCA one protein?
17:42They did not. In fact, when the researchers tried to run the stability calculations on a very specific section of the BRCA1 protein called the RNG domain using the basic alpha fold model, the physics engine returned bad data.
17:55Really? It failed to predict the stability accurately. It failed completely. Why did it fail? If the AI is so good at predicting folds, what tripped it up? Well, it failed because biology is messy. It's crowded, and it's highly collaborative.
18:08The basic AI model generated a prediction of the BRCA one protein sitting completely along in an empty void. Just floating by itself. Right. But in a living cell, the running domain of BRCA1 doesn't work alone.
18:21It relies on forming a complex structure, a header dimer with another protein called BRD one. Furthermore, the 2 proteins are physically locked together by 4 zinc molecules. So the zinc molecules act like biological glue, holding the whole complex together, and the AI didn't know the glue was supposed to be there.
18:40Exactly. The AI lacked the biological context. So the physics engine was trying to calculate the strain on a puzzle that was literally missing half its pieces. Oh, wow. So how did they fix it? The researchers had to manually intervene.
18:53They built a complex model that included both BRCA one and DARD one, and they manually inserted the zinc ions into the simulation. And once they provided the proper biological context, the physics engine's predictions matched reality perfectly.
19:06That is such an important lesson. We can't just feed a sequence into an AI and trust the output blindly. We have to understand the microscopic environment the protein actually lives in. It's a tool, not a replacement for biological understanding.
19:19Right. So, looking at the entire scope of this deep dive. The pattern recognition, the physics engines, the crust versus filling location mapping. How does a clinical geneticist actually implement this tomorrow?
19:32Like, when they're staring at a patient's chart. What do they do differently now? The key takeaway for clinicians is the concept of granularity. Historically, we've relied heavily on single tools that give us broad, sometimes uninformative answers, but this paper demonstrates that clinicians must combine these tools.
19:50So using them together? Yes, by using alphemisms for pattern recognition, running thermodynamic calculations for physical stability, and applying a strict rule about where the mutation lives on the surface.
20:01Clinicians can assign evidence with much greater precision. So they aren't just limited to saying this looks pathogenic or this looks benign. They can confidently declare, we have strong pathogenic evidence, or we have moderate benign evidence, because they have multiple overlapping mechanisms proving the point.
20:17Yes. That granularity is the mechanism that safely pulls those variants out of the VOS limbo. And the clinical recommendation moving forward is that these structural principles shouldn't be limited to just BRCA1.
20:31They need to be systematically tested and integrated into the guidelines for other cancer susceptibility genes as well. To sum up this entire deep dive into the architecture of our biology, how would you capture the core insight of the research teams work?
20:46I'd say the integration of structure informed AI tools, thermodynamic stability calculations, and a crucial understanding of a variant's physical location on a protein, significantly improves our ability to classify misense variants?
20:58This multi-layered approach provides unprecedented evidence, granularity, safely moving more variants out of clinical limbo and into medically actionable categories. It really leaves us with a profound thought to chew on.
21:10You know, what does it mean for the future of personalized medicine when an artificial intelligence's ability to map the microscopic 3D topology of a single protein becomes the deciding factor in whether a patient chooses to undergo a life-altering preventive surgery?
21:24It's just wild to think about. It truly is a new frontier in genomic medicine. Well, this episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
21:39If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
21:53Thanks for listening and join us next time as we explore more science, base by base.