Long-read assemblies and a pangenome reference graph uncover widespread structural variants that shape Mycobacterium tuberculosis evolution and contribute to drug resistance
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Okay, let's impact this. We are diving into a crisis that has really defined global health for centuries, tuberculosis.
0:16Even now, in the age of advanced medicine. This disease, which is caused by mycobacterium tuberculosis or MDB, it just refuses to be beaten easily, just last year. MTV was responsible for 10800000 new cases, and get this, one.2500000 deaths worldwide.
0:33And that devastating statistic, it actually hides an even bigger threat, the one that keeps clinicians awake at night, drug resistance. In 2023, we confirm 400,000 new cases of TV that just defied standard treatment.
0:43These drug resistant strains. They require these complex, expensive and really lengthy therapeutic regimens, which, you know, leads directly to much poorer patient outcomes. Right. And when empty B is drug resistant, our main weapon is whole genome sequencing, we sequence the bacterial DNA, and we try to predict which drug mutations are present.
1:01That tells us which antibiotics still have a shot. But here's the central problem this deep dive confronts. Are we looking in the wrong places? We've become absolute experts at spotting tiny single letter changes in the, uh, single nucleotide polymorphisms or SMPs.
1:17Exactly. For years, the massive, more complex structural changes in the MTB genome. These are changes that can involve deleting or shuffling entire segments of DNA. They've been largely overlooked with all these structural variants or SVs, and they're big enough to knock out a whole gene or drastically change how the bacteria survives inside a person, potentially influencing virulence and resistance in ways we just haven't been able to chart.
1:38And the reason we couldn't see them. Our standard sequencing tools, they simply weren't built for that level of complexity. It's like trying to find a collapsed bridge using a street map that was designed for pedestrians.
1:49It just doesn't work. So this deep dive reveals a revolutionary new mapping approach that finally allows us to lift the veil on these crucial structural variants, fundamentally updating our understanding of MTV evolution and critically diagnosis.
2:04But before we jump into the mechanics, we really want to acknowledge the pioneering work that made this leap possible. Today, we celebrate the team, including Elix Canalda Boltrons, Matthew Silcox, Michael B. Hall, and Sarah J.
2:16Dunstan, along with all their colleagues from the University of Melbourne at the Peter Doherty Institute for Infection and Immunity. Their deep exploration has significantly advanced our understanding of the broader role that structural variation plays in shaping MTB diversity, and most importantly for you listening, how it contributes to the development of drug resistant TB.
2:35To really appreciate the advancement. We 1st need to understand the current baseline. I mean, traditional MTB genomics has always focused on that small scale genetic variability. You know, the SMPs and the small insertions or deletions.
2:49We call them indels, which are usually less than 50 base pairs long. These are pretty straightforward to map to the MTB reference genome. And to be fair, those small changes are crucial. They define the different lineages of MTB, and they make up the bulk of the mutations that the World Health Organization recognizes as causing drug resistance.
3:08So, you know, that's why we focused on them. But the bigger players, these structural variants, SVs, which are defined as anything larger than 49 base pairs, they've always been lurking in the shadows.
3:18And their significance in MTB evolution is undeniable. We know, for instance, that the major TBD one deletion is what separates the older ancestral MTB strains from the modern, globally successful ones.
3:28Right, and we also know SVs can cause drug resistance. We've seen known cases where resistance to drugs like isonizate is caused by a large-scale deletion of the KG gene, or uh, e-thay for ethionomide resistance.
3:43So we know they matter. The barrier has always been getting a comprehensive, reliable list of all of them. And that barrier is purely technological. You think about the sequencing landscape. Most labs globally rely on short read sequencing, predominantly alumina.
3:58Short reads are fast, they're cheap, and they are excellent for finding S&Ps. But when a short read hits a complex region, like a sequence that repeats or a section that is rich in G's and C's. It just gets confused.
4:10And that makes it almost impossible to confidently identify a large insertion or deletion. It's like trying to rebuild a 50 page manuscript using only 2 word fragments. If there's a whole chapter missing.
4:20You can't tell if it's really missing or if it just got moved somewhere else. Exactly. Long read sequencing, like Oxford Nanapore or Pac Bio. It solves this because the reads are long enough to span those complex regions.
4:30But. But they're expensive and they aren't used in routine clinical diagnostics. Right. So the core technical challenge became. How do you reliably extract all of that large SV information using only the short rate data that almost every lad is already generating?
4:45And this is where the genius of the methodology comes in, centered around something called panginome reference graphs or PRGs. So instead of mapping a new MTB sequence against a single linear reference genome, which is like navigating a city with only one predetermined route, a PRG is a complex branched map.
5:02Right. A PRG represents the common sequence that's shared by many isolates, but crucially, it also contains all the known variable sequences, the detours, the deletions, the insertions. They're all integrated right into the map structure.
5:15It's a reference that knows about variation from the very start. And to build this comprehensive blueprint, the team 1st gathered a huge diverse data set. 859 extremely high quality, long read MTV genome assemblies.
5:27This ensures the map isn't biased toward just one single strain. And they were meticulous about getting global diversity. They represented all the major MTB lineages from lineage one all the way through nine, although Lineage 4, which is one of the most successful ones globally, did dominate the cohort, representing almost half the samples.
5:44They use the standard H37 RV reference genome as the backbone for the map, and by incorporating all that variation from the 859 diverse samples, the final MTV BRG. It swelled to contain 900 to 350 different nodes.
6:00It captures this vast genetic landscape. Okay, so building the map is step one. But reading it reliably is step two. Since the existing software they used to build these graphs like minigraph, it didn't have an accurate way to then genotype or, you know, identify the specific SVs present in a new sample.
6:17The team had to invent one. And this necessity led to the creation of mini walk. There's really ingenious piece of software designed to navigate that complex PRG. It identifies all 4 major types of SVs, insertions, deletions, duplications, and inversions by simply tracing the path of a new samples assembly against the fixed reference path within the graph.
6:36It's navigation and comparison at scale. So once they had this new method, the crucial test was, does it actually work better than what we were using before, especially on the short read data, that is the clinicical standard?
6:47Absolutely. And this is the critical benchmark result for you to remember. When they tested the PRG approach using mini walk on standard clinical alumina short red data, it achieved a precision score of .7.
6:59They compared that to manta, which is a leading traditional SV caller based on the linear genome approach, and manta only achieved .46 precision. That is a staggering difference, but let's clarify that for our listener.
7:11What does that jump from .46 to .7 precision actually mean in the clinic? It means reliability. For precision, in this context, is the ratio of true positive SVs found versus all SVs reported. A lower precision score like .46 means that nearly half of the variance man's reports could be false positives.
7:29They aren't actually real. In a clinical setting, calling a false positive SV that looks like it causes drug resistance that could lead you to prescribe the wrong 2nd line drug, and that could be a fatal mistake.
7:40So the higher precision means that the SVs you do find with mini walk are far more likely to be real, giving clinicians much greater confidence in the diagnosis. Precisely. Now there was a trade-off, which you mentioned.
7:51MiniWalk did have slightly lower recall than manta, meaning it missed a few more true variants. But in diagnostics where you need high confidence to make treatment decisions, the gain and reliability for clinically important SVs like large deletions and insertions.
8:05It just far outweighs the risk of missing a few. This firmly establishes the genome graph approach as superior for accurate routine SV detection. With the tool validated, the team could finally characterize the true scale of variation we've been missing, and they cataloged an immense number.
8:213077 unique SVs across the 821 isolates they analyzed. Deletions are the most common with nearly 1700 found, alongside over 1100 insertions. The complexity they uncovered was truly surprising. For example, they found a massive 66.6 kilobase region in lineage 9 that wasn't just deleted or inserted.
8:42It was both inverted and translocated to an entirely different part of the genome. That's a huge genome shaping event that we simply couldn't have mapped accurately before this. And when they mapped these SVs onto the MTB family tree.
8:55The SVs weren't just random clutter. They were actually critical in defining the population structure. They clearly marked the split between the ancient lineages, so L1, L5 through L9. And the more recently evolved, modern lineages, L2 through L4.
9:09And this leads to one of the most interesting evolutionary patterns they discovered, the specialist versus generalist dichotomy. Lineage 4, which is at globally widespread, generalist type, had significantly fewer SVs averaging around 60 to 63 compared to the more geographically restricted, specialized lineages, which average 74 SVs.
9:28That's fascinating. What does that suggest about the role of structural change and success? Are too many SVs, you know, bad for the bacteria, or is there another reason? It could be 2 things. First, maybe too much SV activity hinders global dissemination, making those strains less fit overall.
9:42Second, specialized lineages often undergo what's called reductive genomic evolution. They lose genes they don't need in a specific stable host environment. So the high SV counting specialists might just reflect this rapid shedding of redundant genetic material to optimize survival in a niche.
10:00It also turns out the bacteria aren't making these structural changes randomly across the genome. The break points. The places where the genome is cut or rearranged are heavily clustered in a prominent hotspot between 3.2 and 3.6 megabases.
10:13And that region is famous for being notoriously repetitive and GC rich, containing a dense cluster of PVP genes. These genes are associated with the cell wall and immune interaction. The fact that the SV hotspots are concentrated here suggests these repetitive regions are structurally unstable, making them easy targets for deletion or rearrangement, which the MTB bacterium then uses as raw material for adaptation.
10:35Speaking of adaptation, let's shift to functional systems. The study looked closely at the ESX secretion systems, which are absolutely crucial for MTV virulence. They let it secrete proteins that are necessary for infection.
10:48And crucially, the structural sort of core machinery of ESX1 ESX3 and ESX5 were found to be incredibly conserved, almost 0 SVs impacting them. This just confirms their essential nature. MTB simply cannot survive without them intact.
11:02However, they did find a recurrent frequent deletion in a specific non-core part of the ESX 5 system, the EP 25 PT27 locus. This deletion was found in 9% of all isolates, and critically, it was fixed across the entire sublineage 4.4.
11:19And that matters because this specific deletion is already known to be virulence associated. Prior studies have linked this law to higher bacterial persistence in mouse infection models. So the SE approach confirms this is a consistent, successful evolutionary strategy for MTB.
11:32And in another interesting case of specialization, They found a full deletion of a paralogue, a similar, potentially redundant gene called ESX 5C, and it was found exclusively in the geographically restricted African lineages, L6 and L9.
11:47It seems if you're a specialist in that environment, you can afford to shed that redundant baggage to save energy. But perhaps the strongest signature of positive selection involved genes related to metal homeostasis.
11:58This is a recurring theme in bacterial pathogens. MTB is constantly fighting this war of metal poisoning, having to adapt to shifting levels of zinc, iron, and copper inside the host cell. They screened for SVs that had been acquired repeatedly and independently across the MTB family tree.
12:14This is called homoplessy, and it's basically the genomic fingerprint of strong evolutionary pressure. I mean, why would SVs pop up in the exact same spot in completely different lineages? Because that SV provides a massive survival advantage.
12:27The gene with the strongest signature was RV 3177, which is a probable proxidase, meaning it's involved in detoxifying reactive oxygen species. SVs in this gene were found in 158 isolates across various lineages, suggesting it plays a compensatory role, perhaps helping the bacteria mitigate oxidative stress caused by the host immune system, or by certain antibiotics, like isonyazid.
12:51Another gene that showed this repeating pattern of SV acquisition was CTPG, a zinc exporter. Since MTB rapidly transitions between zinc rich and zinc limited environments in the host, the loss of a functioning zinc exporter might allow the bacteria to hoard zinc when it's scarce, and that would provide a strong selective panage against host defenses.
13:11But the most detailed story of adaptation, though, it's centered on the highly successful lineage one. one sublineage. This drain is globally prevalent, dominating areas like the Philippines, and it's notorious for high transmissibility and drug resistance.
13:23So what SV did they find driving the strains success? They identified specific loss of function deletions of CTPV, a copper exporter. What made this finding so compelling was the clear evolutionary path.
13:35They observed that the deletions grew progressively larger as you moved further away from the ancestral split of the L1.2.1 branch. That indicates continued directional selection for this loss. So the bacteria just kept doubling down on shedding this copper exporter.
13:51Why? I mean, copper is essential for us, but the host immune system uses copper to try and poison the bacteria inside the macrophage. Exactly. When they perform transcriptoma comparisons of an L1.. isolate with this deletion against other L1 strains, it suggested a state of copper accumulation inside the cell.
14:08They saw other copper associated genes, like CTPG being enriched. This SV driven loss of the exporter essentially provides the bacteria with increased copper tolerance, allows it to survive better in the toxic copper rich environment of the host granuloma, contributing directly to its higher virulence and success.
14:25The ultimate application of this research lies in diagnostics. This is where the PRG approach scales up to become a true game changer. They didn't stop at 859 assemblies. They use the PRG and mini walk to screen over 41,000 existing clinical shortread isolates for drug resistance associations against 22 different drugs.
14:44And the ability to mine that massive data set with high confidence immediately provided new resistance markers we'd never seen before. They cataloged 14 non-canonical SVs and 27 SV affected genes that were associated with resistance across 10 different drugs.
14:59The findings for isoniazid were particularly striking. They identified a deletion overlapping a gene called RV 3434C that was present in 99% of resistant isolates. This just confirms the clinical relevance of that specific structural change.
15:13But here is the major clinical takeaway, the one that justifies this entire technical shift. They found resistant isolates with the specific SVs that lacked any known resistance mutations listed by the WHO, or standard diagnostics like TB Profiler.
15:28This proves that structural variants are an entirely hidden layer of resistance mechanisms, currently being missed by existing standard sequencing workflows. If you are missing the SV, you are missing the resistance.
15:39And the relevance spans the entire treatment spectrum. SVs were linked to resistance to traditional drugs like PZA and EMB, and critically, to key 2nd line and new novel drugs like Amaxen, Badok line, BDQ, Capriomycin, and Delamened, or DLM.
15:56Furthermore, the study confirmed that SVs are not just messing with the coding regions, they affect non-coding or intergenic regions too, which underscores the necessity of sequencing and accurately analyzing the entire genome.
16:07We also saw low predictive importance scores for the newer drugs like BDQ and DLM, which only reinforces the urgent need to find new resistance variants, like these SVs to improve the efficacy of these crucial last resort treatments.
16:19So this graph-based approach isn't just an academic curiosity. It's an ideal candidate for integration into existing clinical whole genome sequencing pipelines like right now. Well, the advancement is massive.
16:30The researchers were clear about the immediate next steps. Methodologically, the current mini walk tool requires preassembled genomes. To truly boost recall and speed, future development needs to allow it to genotype SVs directly from unassembled short reads.
16:45And biologically, we need to formally confirm that L1.2.1 copper story. The hypothesis that the CTPV deletion causes copper accumulation is strong, but we need direct validation. We need intracellular copper quantification.
16:59And crucially, we need animal infection models that replicate the high copper levels found in human granulomitous lesions to confirm that this deletion truly provides an adaptive advantage in Vivo. To synthesize the core message of this deep dive.
17:11Structural variants in MDB are not mere genetic noise. They are critical recurring evolutionary forces that drive adaptation, especially in virulence and drug resistance. The development of panginome reference graphs, and high precision tools like mini walk, provides the high-confidence framework necessary to systematically map these crucial large scale changes for the very 1st time.
17:32This technological shift from outdated linear-based sequencing diagnostics to panginome-based diagnostics is essential if we hope to capture those hidden resistance mechanisms, especially for widely used and key 2nd line drugs where the stakes for patient survival are highest.
17:48So, as we successfully integrate SV detection into routine diagnostics, what does this essential new genetic layer mean for designing more targeted personalized therapeutic strategies against dominant and highly transmissible specialist lineages, such as the successful L1.2.one sublineage?
18:06This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating.
18:19If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
18:28Thanks for listening and join us next time as we explore more science, base by base.