This episode explores a multi-omic study showing that mass spectrometry–based chemoproteomic detection of cysteine, lysine, and tyrosine (CpDAAs) highlights protein sites and regions enriched for pathogenic missense variants and variant uncertainty.
0:00Welcome to Base by Base, the paper cast that brings genomics to you, wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Now, imagine you're facing, um, a major real world clinical dilemma.
0:14You've been dealing with some severe, mysterious health issues. Right, like something totally debilitating. Exactly. So your doctor orders a comprehensive genetic test, hoping to finally pinpoint the cause.
0:24Weeks pass, you know, you're waiting, the results come back. But instead of giving you a clear answer. The doctor just points to a specific spot on your DNA and says, well, it's a variant of uncertain significance.
0:35Ah, V-U-S. The dreaded VUS. Yeah, of US. Basically, they found a typo in your genetic instruction manual, but absolutely nobody knows if it's completely harmless, or if it is, like, the exact root cause of your desires.
0:49It's just the ultimate diagnostic guessing game, honestly, for doctors and patients alike. I mean, we have sequenced 1000000s of human genomes over the last couple of decades. Yeah, mountains. Right? Mountains.
1:00But our ability to actually interpret what those sequences do, how the single letter change affects a living, breathing human, has severely lagged behind our ability to just, you know, read the code. So for today's deep dive, we've pulled together some fascinating research to help you understand a potential way out of this limbo.
1:17How could this change if we stop just looking at the flat DNA sequence on a screen and start looking at how the resulting protein chemically reacts in the real world. Yeah, it's a great question. Right.
1:29What really happens when we combine massive genomic databases with actual protein chemistry to finally solve this diagnostic guessing game? I mean, it represents a complete paragigm shift. Instead of guessing a mutations impact based purely on the genetic text.
1:44Researchers are now mapping the active physical battlegrounds on the proteins themselves, you know, to see where a typo actually causes structural damage. Today we celebrate the work of Maria F. Palifox, Carrion and Bacchus, Valerie A.
1:56Arbeleta, and their team at UCLA, who have advanced our understanding of prioritizing disease associated genetic variants. Yeah, their paper is titled prioritizing disease associated misins variants with chemoproteomic detected amino acids.
2:10And that was published in the American Journal of Human Genetics, volume 112, on July 3, 2025. This research tackles one of the biggest, just most persistent hurdles in modern clinical medicine, the missends variant.
2:23Right. And to give you some context on why this matters. A mis sense variant is actually the most common type of protein altering genetic variation we see in humans. It's just a single letter swap, right?
2:33Exactly. It happens when just a single letter in your DNA code is swapped out. Because DNA is read in triplets to build proteins, that one wrong letter results in the substitution of just one single amino acid for another in the final protein chain.
2:46Okay, let's unpack this. It's just, you know, one single building block being swapped out of a protein that might be 100s or 1000s of blocks long. And the scale of the confusion this causes is just immense.
2:58If you look at Klinvar, which is, um, the massive public database where clinical labs upload genetic variants, they find in patience, over half of all misense variants reported are classified as variants of uncertain significance.
3:12Over half. Wow. Yeah, more than 50%. And we know these variants are incredibly important. I mean, misance variants are responsible for causing roughly 38% of all men dealing in disorders. These are the single gene diseases, right?
3:24cystic fibrosis or sickle cell anemia. Exactly. But finding the truly dangerous mutations buried inside this mountain of harmless ones is a massive challenge. A single missence variant might do absolutely nothing to you.
3:38Or it might completely misfold a crucial protein and cause a devastating life-threatening disease. Let's use an analogy here to see if I'm tracking. If we think about your DNA as a giant, incredibly dense instruction manual for building a car engine.
3:52A misins variant is essentially knowing that a single word was changed somewhere on like page 400. You know the word is different, but you need to figure out whether the change is a minor detail, like changing the instruction from paint the cap blue, to paint the cap red, or if it's a crucial operating step, like remove the screw instead of tighten the screw.
4:12It's a perfect way to look at it. And to take that engine analogy a step further. The problem is that our current tools for figuring out whether that word is dangerous are essentially just reading the manual over and over again.
4:24Right now, clinics rely heavily on predictive computer models. Um, like the KED score to guess whether that changed word will break the engine. I read a bit about CAD scores in the background material.
4:36They rely heavily on evolutionary conservation, right? Like looking back in time. Yes, exactly. A kid's core looks at a specific gene and asks, has this particular amino acid changed at all over 1000000s of years of evolution?
4:48It compares the human gene to the same gene in a mouse, a dog, or, you know, a fruit fly. If an amino acid has remained exactly the same across all those species for 1000000s of years, the computer assumes it must be incredibly important.
5:02So, if a human patient has a mutation there, the computer flags it as likely dangerous. But wait, so CAD just looks at evolutionary history. It doesn't actually test the physical protein in a biological environment to see if it's broken.
5:16No. And that is the fundamental limitation. These tools are incredibly useful as a 1st pass, sure. But they lack chemical precision. They are still just making an educated guess based on the text of the manual.
5:29Right. And for a patient sitting in a clinic awaiting a definitive diagnosis for a severe illness. I mean, an educated guest simply isn't enough. They need physical proof. Which brings us to the core methodology of this UCLA team.
5:42They didn't just look at the code. They looked at the chemistry. They used an approach called chemoproteomics. Yeah, a bit of a mouthful. Definitely. Now, that is a dense word. I know proteomix means studying all the proteins in a cell, but what is the chemo part doing here?
5:55How do you actually physically probe a microscopic protein to see what's happening? Are they like looking through an ultra-powerful microscope? No, it's actually much more dynamic than a microscope. Think of chemoproteomics, like throwing a specific type of chemical paint at an invisible wall to see what sticks.
6:10Okay. The team uses chemical probes, these tiny, highly reactive molecules and drops them into a soup of human cells. These probes are designed to seek out and bind only to parts of protein that are highly chemically reactive.
6:24Oh so they are fishing for hotspots. They are. fascinating here is they specifically focused on 3 amino acids that are known to be potentially reactive, cystine lycine and tiracine. Okay, CK and Y. Exactly.
6:37When the probes find a highly reactive version of one of these amino acids, they permanently attached to it. Then the scientists use mass spectrometry. It is essentially a highly precise biological wing scale, to identify exactly which proteins and which specific amino acids on those proteins caught the probe.
6:56Right, what got heavier. Exactly. And the team calls these highlighted spots. Chemoproteomic detected amino acids or CPDAAs. Okay, so they map out all these sticky, highly reactive hotspots across the human prodium.
7:10Then, according to the paper, they cross reference those physical coordinates against two massive databases. OMIM, which catalogs human monogenic disease genes, and Clinvar, which tracks known pathogenic variants.
7:23Yep, bridging the chemistry in the genetics. But I have a fundamental question about the biological logic here. Why does a highly reactive amino acid automatically Sigroll at a specific spot on a protein is functionally important?
7:34Like, why should we care if it's sticky? Well, if an amino acid is highly reactive in a biological environment, it means it is chemically primed and exposed. It is ready to interact with other molecules to buy into a substrate or to catalyze a crucial chemical reaction.
7:49It's out in the open doing the heavy lifting for the protein's function. Because it's actively working, it's vulnerable. Vulnerable is the key word. It is a critical node. If a genetic mutation swaps out that specific, highly reactive amino acid for something dull and unreactive, you are effectively breaking the exact tool the protein uses to do its job.
8:11Going back to the engine, it's like finding the gears that are actually coded in oil, actively generating friction and turning the belt versus, you know, the decorative plastic casing that just sits there.
8:20If you snap the turning gear, the car stops. That's spot on. And when the team mapped these reactive hotspots against the genetic databases, The initial gene level findings completely validated that logic.
8:31They discovered that proteins containing these reactive CPBAA sites are highly enriched for known monogetic disease genes. And interestingly, they are highly enriched for FDA approved drug targets. Furthermore, these genes are generally what geneticists call highly constrained.
8:49Meaning the human population rarely tolerates mutations in them. The evolutionary filter weeds them out because altering these genes is so damaging to human survival. But wait, let me push back on that for a second.
9:01If evolution weeds these damaging mutations out. Why are we seeing so many of them popping up in the Klinbar database causing human diseases today? It's a great point. Evolution is a slow population level filter, but biology is constantly mutating.
9:17Many of the severe diseases we see in Glenvar are caused by de Novo mutations. Spontaneous ones? Yeah, spontaneous new typos that happen during sperm or egg formation, so the parents don't have the disease, but the child does.
9:28Or they're recessive traits, meaning a person can carry one broken copy of the gene without symptoms, hiding it from evolutionary pressure until 2 carriers have a child. Ah, I see So the mutations are constantly appearing in the clinic, even if they don't spread widely through the evolutionary tree.
9:44That makes sense. So the team establishes that these reactive sites are on important genes, but they didn't just stop at the gene level, they zoomed in to look at the spatial relationships, how close the known disease causing mutations were to these reactive hotspots.
9:58They did. And they looked at this in two distinct ways, right? In one D and 3D. By one D, I assume you mean looking at the linear sequence of the protein, like stretching out a string of yarn. Right. When they looked at the one delinear sequence, they found that pathogenic variants were significantly enriched within a very short distance, just 6 amino acids away from a CPDAA hotspot.
10:19That's super close. It is. But as we know, proteins don't operate inside your body as flat stretched out strings of yarn. They fold up into incredibly complex, three dimensional origami structures. So how did they analyze the 3D space?
10:34Because that seems much harder to calculate than just counting letters in a line. It is. And they had to use advanced protein structure databases to do it. When they look at the proteins in their fully folded 3D shapes, the results were striking.
10:46They found that pathogenic variants clustered incredibly close to these reactive sites, specifically within an 8 angstrom radius. Okay, 8 angstroms. Without a background in structural biology, that number doesn't mean much to me.
10:58Give me a sense of scale. Well, an Angstrom is one 10 billionth of a meter. Okay, tiny. Super tiny. To put an 8 angstrom radius in perspective. We were talking about a distance roughly the width of just 3 or 4 individual atoms.
11:12It is a tiny, incredibly localized micro environment. I see. If we use a crumpled paper analogy, like if you take a flat piece of paper and draw a red dot at the top and a blue dot at the bottom, they are very far apart in one D space.
11:26But if you crumple that paper into a tight ball, suddenly the red dot and the blue dot might be physically touching each other. And that's exactly what is happening in the protein. Even if a genetic mutation happens far away on the linear sequence.
11:39If the protein folds in a way that brings that mutation right into the 8 angstrom physical space of the reactive hotspot. It is highly likely to cause disease. Oh, because it interferes with the hotspot.
11:51Exactly. The mutation might physically block the reactive site, alter its shape, or change the local electric charge so the reactive amino acid can no longer do its job. The researchers also broke down the differences between the 3 amino acids they studied.
12:04I noticed that lycine and tyracine sites showed a very strong association with functional significance and overlapped heavily with these pathogenic variants. They did, yeah. But I'm looking at your notes here, and there's a really counterintuitive detail about Sistine.
12:18If Sistine is this incredibly reactive hotspot, Why did the study find it is actually depleted, meaning there's less of it than expected in these disease genes overall. It seems like a paradox at 1st glance, but it actually highlights the unique chemical properties of Sistine.
12:33You're right, globally across these monogenic disease genes. Sistine is depleted. But the researchers point out that while it is rare, when a mutation causes the loss of a sustine, that mutation is massively enriched for causing disease.
12:48Wait, why the discrepancy? Because in human biology, many cystines are involved in forming what are called disulfide bonds. You can think of these as strong permanent steel beams that hold the protein's 3D origami shaped together.
13:02They are load bearing structural pillars. Okay I can picture that. Because they are locked into these strong structural bonds, they aren't freely reactive. Oh, so the chemical paint from the chemoproteomic probes just bounces off them.
13:15Precisely. Standard chemocodeomic experiments often fail to detect these specific structural cystines because they are already chemically occupied holding the protein together. That makes total sense. But the cystines they did manage to detect, the ones that are not structural, but are exposed, free floating, and highly reactive.
13:31Those are absolute beacons for disease when they're mutated, or when mutations happen within that 8 angstrom radius. That distinction sets up the real world validation perfectly. The team didn't just leave this as an abstract computational exercise matching databases, right?
13:47They actually took this theory into the lab to prove it in living cells. Yes they did. They focused on a specific case study, the fumerite hydratase enzyme, or FH. Yeah, FH is a critical metabolic enzyme.
13:59When it malfunctions, it is linked to a severe hereditary form of kidney cancer, as well as other devastating metabolic conditions. Terrible. It is. Now, for FH to function properly, it has to assemble into a Tetramer, meaning 4 identical copies of the protein have to physically click together into a complex machine.
14:17The team zeroed in on a specific, highly reactive Sistine hotspot on this enzyme known as C33. So applying their new map to the specific enzyme, did the mutations cluster around C 333 like the data predicted?
14:30The statistics on this single hotspot were staggering. When they mapped known genetic variants onto the 3D structure of the FH enzyme, they found that within an 8 Angstrom sphere of this one reactive C's 333 site, 77% of the surrounding genetic variants were classified as pathogenic or VUS.
14:5077%. That is a massive cluster of danger right at that physical location. It is the absolute definition of a functional hotspot. And to prove that this proximity actually causes the disease. The team experimentally introduced the surrounding misence variants into cells in the lab.
15:07To see what would happen. Exactly. They show that mutating the areas immediately around this reactive site, physically warped the enzyme. It broke the protein's ability to assemble to the four-part tetramer.
15:18The machine literally fell apart, which, in the human patient, leads to the toxic buildup of metabolites that drives the cancer. So what does this all mean? We now have this incredible multi-omic map that overlays the genetic code with physical, chemical vulnerabilities.
15:32How does this change the landscape of rare disease treatment and diagnosis moving forward? Well, the immediate clinical value here is immense. We can now use this chemoproteomic detection synergistically with those predictive computer tools we discussed earlier, like the CAD score.
15:46Combining them. Yes. Imagine a patient gets a VUS diagnosis today. The can score says, well, the evolutionary text looks suspicious. It's rarely changed in mice. But the doctor still isn't sure. Now, you overlay this new chemical map.
16:01If that mysterious VUS is sitting within 3 to 4 atoms 8 angstroms of a highly reactive CPDAA hotspot, the clinical geneticist can confidently upgrade that VUS into a clearly understood pathogenic variant.
16:16It speeds up the diagnosis tremendously. Gives doctors the hard physical evidence they need to look a patient in the eye and say, yes, this specific typo is the cause of your disease. Right. It gives them certainty.
16:27That ends the exhausting diagnostic odyssey that so many rare disease patients go through. But reading the paper, there is also a huge implication for drug development here, isn't there? Absolutely. The paper highlighted that only about 13rd of the proteins associated with rare disorders actually overlap with the specific reactive sites they mapped.
16:44So it's not a silver bullet for every single genetic disease. But for that one third that do overlap. What does that mean for pharmacology? For those that do overlap, these highly reactive pockets offer the perfect targets for new therapies specifically covalent drugs?
17:01I've heard that term covalent drugs. How is that different from, say, popping an ibuprofen for a headache? Most standard drugs, like ibuprofen, bind to a protein temporarily. They attach, do their job, and eventually float away.
17:16A cogolent drug is different. It forms a permanent, irreversible chemical bond with its target. Oh, wow. Yeah. Aspirin is actually a classic example of a covalent drug. Because these CPDAA hotspots are already naturally highly reactive, they are biologically primed to form these permanent bonds.
17:33Oh, I see. So if a pharmaceutical researcher is trying to figure out how to drug a misbehaving protein, they don't have to guess where to aim. This map gives them a literal bull's eye. A structural bullseye.
17:44They can design a covalent compound to permanently latch onto that exact reactive pocket to fix or inhibit the protein. Exactly. And considering less than 5% of rare diseases currently have FDA approved drugs, presiding researchers with pre-validated, highly reactive structural targets is a massive leap forward.
18:03It's the bridge between reading the flat genome and actually drugging the 3D protium. Of course, as with all groundbreaking science, we have to talk about the limitations. Where does this approach still fall short?
18:15Well, the chemoproteomics tech we have right now currently misses many sites in the human body. As we discuss with cystine, if a reactive site is buried deep inside a protein or locked up in a structural bond, the chemical probes can't reach it.
18:28They just bounce off. Right. Furthermore, the Clinvar database itself isn't perfect. It is crowdsourced from 1000s of clinical labs around the world, and you often have conflicting interpretations of the same genetic variant depending on which lab uploaded the data.
18:42The output map is only as good as the input data. So what's the next logical step to clear up that messy data? The next frontier is combining this chemical reactivity mapping with CRISPR gene editing technology.
18:54Right now, we map the hotspot. The next step is using CRISPR to systematically introduce every possible mutation into that 8 angstrom micro environment in living cells, testing the exact severity of every single typo in a high throughput way.
19:08So to synthesize all of this, By integrating flat genomic variant data with the 3D physical chemical reactivity of proteins, scientists can finally pinpoint the functional hot spots that drive human disease.
19:21Yep, that's it, exactly. This multi-omic map transforms how we prioritize mysterious mutations, bringing us closer to ending genetic guessing games and designing highly targeted, permanent, covalent drugs.
19:31It moves clinical genetics from reading a two-dimensional sequence to mastering the three-dimensional chemically active reality of human health. Which leaves you wondering about the future. If we can eventually map the exact reactive vulnerabilities in 8 angst microenvironments of every single protein in the human body, how long until artificial intelligence can bypass clinical drug trials entirely, simulating these physical chemical interactions perfectly in a computer to cure diseases we haven't even named yet.
19:59What does this mean for the 1000000s of patients trapped in genetic limbo and how fast can we turn these reactive hotspots into life-saving therapies? That's the $1000000 question. This episode was based on an open access article under the CCBY 4.0 license.
20:14You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
20:27Now stay with us for an original track created especially for this episode, and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science, base by base.