This review shows how short-read shotgun sequencing and genome skimming recover organellar genomes, estimate genome size and repeat content, and enable scalable biodiversity monitoring.
0:00Welcome to Base by Base, the paper cast that brings genomics to you, wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. We have a really fascinating mission ahead of us today.
0:11We absolutely do because, you know, to understand why our topic today matters to you. We have to look at, well, we have to look at a massive global challenge we are currently facing. Right. The coming Montreal global biodiversity framework.
0:24Exactly. Often just called the GBF. It is this monumental international initiative. And the primary goal is to halt and reverse human induced species extinction by the year 2030. Which is basically right around the corner.
0:38It really is. And the framework is designed to ensure the sustainable use of biodiversity, protect terrestrial and marine habitats, control invasive species. All nobble goals. All of it, restore degraded ecosystems globally.
0:51But there is a massive glaring roadblock standing in the way of all those targets. We can't save species we haven't even identified yet. Exactly. We cannot accurately monitor complex ecosystems if we don't know the foundational building blocks making them up.
1:05And to achieve those ambitious GBF targets. The scientific community desperately needs robust, high resolution monitoring systems. I mean, we require comprehensive knowledge of individual species. They're geographic distribution, their population genetics.
1:22Which traditional taxonomy just can't keep up with. No, not at all. Traditional taxonomic approaches. Well, absolutely foundational to the history of biology. They're simply too slow. They are too labor intensive when rapid biodiversity assessments are needed in the face of, you know, accelerating environmental change.
1:39We need genomic data. Right. And we need to generate it on a massive global industrialized scale. So that brings us to the core of today's deep dive. We're jumping right into a compelling paver from early 2026, published in the journal Trends in Genetics.
1:53The title is The Untapped Potential of Short Red Sequencing and Biodiversity Research. And let me tell you, this one really flips the script on how we typically think about cutting edge science. Today, we celebrate the work of the authors and researchers behind this paper.
2:06who have advanced our understanding of global conservation technology. Because this paper fundamentally challenges this innate bias, we tend to carry. The assumption that whatever is newest, flashiest, and most expensive, is inherently the best tool for every job.
2:22Right. This paper takes a hard look at the reality of global conservation, and argues the exact opposite. Okay, let's unpack this because in the world of DNA sequencing right now. The entire scientific community seems obsessed with long read sequences.
2:36Whoa, absolutely. the shiny new toy in the lab. But this paper argues that older, vastly cheaper technique, known as short red or shotgun sequencing, is actually the rugged, versatile secret weapon that is going to successfully map Earth's biodiversity.
2:51And to be fair, long read technology is undeniably fantastic for specific applications. I mean, it produces continuous reads of thousands, sometimes even 1000000s of base pairs in a single stretch. Which is amazing.
3:03It is an amazing capability. If your goal is assembling a pristine gold standard reference genome from scratch. But out in the messy real world, short read sequencing is far more universally applicable.
3:16It's highly scalable. Exactly. It provides an easily generated universal data source that seamlessly spans all biological levels, from the genome of a single individual, up to the complex genetic soup of an entire ecosystem.
3:30The technology has just been quietly driving down costs and enhancing performance in the background to the point where it is now uniquely positioned to help us actually achieve those GBF objectives. It's the workhorse we need right now.
3:43So to really help you visualize the mechanical difference between these 2 technologies. Think of an organism's DNA, like a massive, thick encyclopedic book. Okay, a big book. Right. Long read sequencing is like someone handing you intact, beautifully preserved chapters of that book.
3:58It is relatively easy to put the narrative of the story together. Short read sequencing, on the other hand, is like taking that entire encyclopedia and running it straight through an industrial paper shredder.
4:07A complete mess. You are left with a massive pile of 1000000s of tiny snippets. In the case of short red sequencing, those shredded DNA fragments are typically only 50 to 300 base pairs long. Very small shreds.
4:21Right. And then you have to use heavy duty bioinformatics to computationally piece those fragments back together based on overlapping sequences. And what's fascinating here is that while the paper shredder approach sounds chaotic.
4:33And unnecessarily difficult compared to the chapter by chapter method. Out in the real world, nature has already run the book through the Shredder for us. Right. Nature doesn't gently hand us pristine, highly preserved chapters of DNA when we are trudging through a rainforest or a desert.
4:48No, the natural environment is incredibly hostile to genetic material. Out in the field. DNA degrades with astonishing speed. It just falls apart. Exactly. It is constantly subjected to enzymatic digestion, intense ultraviolet radiation from the sun, extreme temperature fluctuations, oxidation, hydrolysis, and aggressive microbial activity.
5:10All of these environmental factors physically break down the long DNA strands. And they chemically modify the individual bases too. So long read technology is... Look, underweight, perfectly intact strands of DNA to function properly.
5:28So if you feed it the degraded DNA found in nature. The sequencing process simply fails. But short read sequencing is fundamentally designed from the ground up to work with fragments. It thrives on those naturally shredded pieces.
5:40Which perfectly transitions us into one of the most exciting applications of this technology. A field the paper refers to as museum mix. Museumics, yes. If you think about it, museum collections hold the ultimate time machine for biodiversity.
5:53We are talking about vast, dusty archives filled with extinct, incredibly rare, and entirely inaccessible species that you could never find by doing field work today. Natural history museums around the world function as these massive untapped biobanks.
6:08But the specimens housed inside them present a unique challenge. Because they're old. Right. Whether we were talking about subsossil remains, insects that have been dried and pinned to boards for a century, or tissue samples preserved in jars of formalin.
6:22They all contain what we call historical DNA or ancient DNA. And this genetic material is highly fragmented? heavily chemically altered by the preservation methods themselves. So before the recent advances in short read sequencing, these historical specimens were essentially locked away from modern genomic analysis.
6:41They were invaluable for studying physical morphology. Looking at their shapes and structures. But entirely genomically inaccessible. But now, precisely because short red technology loves those tiny chemical shreds, we are finally unlocking those archives.
6:55And the scale of the examples provided in this paper is just staggering. The authors highlight a specific initiative at the Australian National Insect collection called the Barcode Blitz. It's great example.
7:05In a period of just 6 weeks, they managed to digitize in barcode 41,000 Lepidopter specimens. That is 41,000 individual moths and butterflies. Presenting over 12,000 distinct species. They achieved an 86% success rate from dry pinned insects.
7:22Incredible throughput. And the incredibly degraded DNA extracted during that initiative can now be utilized for massive, large scale, short red genomic data sets. It represents an unprecedented scaling up of reference sequence generation that simply wasn't possible a decade ago.
7:39And the paper doesn't just stop at insects. There was another massive study mentioned where researchers successfully mapped 181 hapless glared demasponge specimens. Which is a huge deal. Right. For you listening, Demispunges are notoriously difficult to sequence because they are basically the living water filters of the ocean.
7:57They just suck everything. They are absolutely packed with environmental contaminants, symbiotic bacteria and degraded matter. To pull cleave genetic data from historical sponge specimens is a monumental bioinformatic triumph.
8:10If we connect this to the bigger picture. Sequencing these specific historical collections introduces the critical concept of type genomics. Type genomics Right. In the rigid rules of biological taxonomy, a type specimen, specifically a holotype, is the exact physical museum specimen upon which the scientific name of an entire species is officially based.
8:33It is the absolute gold standard for identification. Because right now, if a scientist pulls a random sequence from a public database, there is a decent chance it might be misidentified or contaminated.
8:45Which throws off all future research based on it. The databases are, unfortunately, riddled with errors from decades of varying technological standards. So if we want to monitor global ecosystems accurately, we desperately need pristine, validated reference databases to compare our newly collected environmental samples against.
9:03Exactly. By extracting the degraded DNA directly from the actual original name bearing types, specimens in museum vaults, and using short read sequencing to generate validated genomic reference sets, we are permanently cleaning up the databases for all future scientific endeavors.
9:17Here's where it gets really interesting. Because the paper shifts focus from these individual, carefully pinned specimens in museum jars to looking at entire chaotic environments using a technique called genome skimming.
9:32Genome skimming. When I 1st read that term, genome skimming, it sounded a bit illicit, like an accountant skimming off the top of the books. A little sketchy. But in genetics, it refers to intentional, extremely low coverage sequencing.
9:46In traditional sequencing, if you want to understand an entire nuclear genome, you have to sequence it deeply. You might need a coverage of 30 or 40 times the genome's total size just to ensure every single base pair is red accurately and overlapping correctly.
10:00That sounds expensive. It is. But genome skimming throws that requirement out the window. It intentionally targets a sequencing coverage of only 5 x or sometimes even an ultra low coverage of less than one x.
10:12Which initially sounds totally counterintuitive. How does writing less than one full copy of a genome give you any useful information at all? The secret lies in cellular architecture. You aren't actually trying to reconstruct the entire massive nuclear genome.
10:26You are actively and intentionally targeting only the most abundant pieces of DNA within the cell. Ah okay. Think about it. A cell only has one nucleus. But it might have 100s or 1000s of mitochondria.
10:40Plants have 1000s of chloroplasts. Exactly. There are also specific nuclear ribosomal repeats that occur in massive numbers. Because these specific genetic elements exist in such high copy numbers per sell, even a very light ultra low skim of the total DNA guarantees that you will randomly hit and capture those high copy targets.
11:00It is basically sweeping up the most common genetic dust in an environment. sweeping up the dust, yes. And for you listening, try to imagine the practical power of this out in the field. Researchers no longer have to spend weeks trekking through a jumble trying to catch one specific rare butterfly.
11:14They can just set up a malaise trap. Right. Which is essentially a specialized tent that funnels all flying insects in an area into a single collecting bottle, creating a massive bulk sample. Or they don't even need the insects themselves.
11:25They can just scoop up a handful of dirt or take a liter of ocean water and analyze the environmental DNA floating in that genetic soup. This approach represents a massive fundamental upgrade from how the scientific community used to analyze those complex bulk samples.
11:43Previously, the dominant method for EDNA was metabar coding, which relied heavily on a technique called PCR or polymerous chain reaction. Right, to amplify a single standardized barcode region of DNA so it could be read.
11:57But the paper explicitly points out that relying on PCR amplification has a fatal flaw when it comes to environmental monitoring. Primer bias. Primer bias. The chemical primer is required to kickstart the PCR amplification process might accidentally bind perfectly to the DNA of one specific species of beetle in your soil sample, but bind very poorly to a different species of beetle right next to it.
12:19So when you look at the final data. The results are heavily skewed. It makes it look like the 1st beetle is incredibly abundant and the 2nd beetle barely exists, entirely because of a chemical quirk in the test.
12:29And that primer bias is the exact reason environmental monitoring has been so frustrating historically. You simply could not use PCR metabar coding to accurately estimate the actual biomass of the species in an ecosystem.
12:43But genome skimming. often applied as mitochondrial metagenomics. Completely bypasses the PCR staff. Yes. You extract the total DNA from your bucket of ocean water or soil, run it directly through the short ridge shredder, and sequence it exactly as is without any artificial amplification.
13:00The paper notes this drastically improves the correlation between the read count in the computer and the actual living biomass in the environment. Scientists are now using this to accurately estimate the true biomass of marine macrobenthos, you know, the complex communities of creatures living on the ocean floor from single book samples.
13:18Now, reading this paper, the computational side of that process seemed like an impossible hurdle. It does sound overwhelming. You take a bucket of ocean water containing the DNA of 10,000 different organisms, run it all through the short red paper shredder, and you end up with 1000000000s of tiny 150 base pair shreds completely mixed together.
13:38Massive digital jigsaw puzzle. Exactly. But the bioinformatic workaround they describe is brilliant because they don't even try to put the whole puzzle back together. They skip the assembly process completely.
13:49That is the true bioinformatic magic that makes his entire global endeavor possible. We don't need to assemble the genomes to understand exactly how these species are related or what is present in the sample.
14:00The field relies heavily on what are known as assembly free methods, utilizing advanced software tools with names like squammer, mash, and reed 2 tree. And one of the core approaches these tools use is hunting for USCOs.
14:14Universal single copy orthologs. Yes. If we go back to our shredded encyclopedia analogy. This isn't trying to reconstruct every page. This is like writing a software program that just sifts through 1000000s of paper shreds looking for one specific, highly unique sentence that we know exists in almost every book.
14:30Like the copyright phrasing. That is a highly accurate analogy. USCOs are highly conserved housekeeping genes. They are the genetic instructions for essential, basic cellular functions that almost every living organism shares, and they typically only appear as a single copy in a genome.
14:48So instead of wasting massive supercomputer resources, trying to assemble the entire genome from scratch. The software rapidly sifts through the unassembled shreds. identifies the fragments containing these specific housekeeping genes, and uses those isolated fragments to map the evolutionary relationships between the organisms in the sample.
15:07The paper also details another assembly free method that sounds even more abstract, involving something called kamers. Ah, yes, Kamers. Instead looking for specific genes, they're just looking at mathematical patterns in the freads.
15:19Right. Cameras are simply short, contiguous sequences of DNA of a predefined length. For example, a string of exactly 21 base pairs. The bioinformatic software bypasses traditional sequence alignment entirely.
15:31It doesn't compare the shreds to a reference picture. Instead, it mathematically analyzes the frequencies and distributions of these 21 letter patterns directly from the raw, unassembled data pile. So by statistically comparing the mathematical distribution of these comber patterns between 2 different samples, the software can accurately estimate the evolutionary distance between them.
15:52It is a brilliant mathematical shortcut that saves an immense amount of computational power in time. And it allows us to map the evolutionary tree of life faster than ever before. But moving on, I have to point out one of my absolute favorite paradoxes from this paper.
16:07The contamination aspect. Yes. In traditional old school genomics, if you were sequencing a rare butterfly and your sample got contaminated with bacteria or a common fungus, you threw the sample away. ruined your day.
16:19Your data was considered useless, but with modern short rid sequencing, this broad contamination is suddenly treated as a massive feature, not a bug. This shift in perspective, revolves around the concept of the hollow bind.
16:32Biologists increasingly recognize that an organism does not exist in isolation. A hollow viant is the host organism combined with all of its microbial hitchhikers. The specific bacteria, fungi, and viruses that live on its surface and inside its gut.
16:47Together, their combined genetic repertoire forms what we call the hollogenome, which heavily influences the host's survival and evolution. Because when you extract the DNA from a whole insect, you aren't just getting the insect's DNA, you are getting the DNA of everything it recently ate and every microbe living inside it.
17:06And the short rid sequencer just blindly shreds, and reads all of it together. The paper cites a truly fantastic example of this, researchers were able to recover the complete high quality genomes of Wabakia symbians directly from the short red sequencing data of their insect hosts.
17:22Wobakia is a highly influential bacteria that manipulates the reproduction of its hosts. And the researchers didn't have to isolate the bacteria or try to culture it in a petri dish in a app. They simply computationally separated the bacterial genome from the host's genetic shreds after the fact.
17:38So we can essentially study the host and its entire microbiome at the exact same time from the same sample. But the researchers are finding even more secrets hiding in that low coverage genomic dust. Using specialized software tools like respect and modest, they are actually estimating the total size of an organism's genome and profiling what they call the Mobilome.
18:00The mobile loam, which sounds like a fleet of vehicles, but it actually refers to transposable elements or jumping genes. For a very long time in genetics. Transposable elements were dismissed as selfish genomic parasites.
18:13They were considered junk DNA that simply used the host cellular machinery to constantly copy and paste themselves into different locations across the genome. Just clogging up the works. Right. But modern genomics reveals that this activity actually drives massive genomic variation and can promote major evolutionary innovations.
18:30Even from highly fragmented, low coverage short read data. We can now track the abundance of these jumping genes and identify sudden bursts of transposition activity that correlate with major evolutionary shifts.
18:43It is incredible what we can extract from this data now. So what does this all mean for the immediate future of conservation? We have explored this incredibly versatile tool that can read degraded historical DNA from a century ago, accurately estimate the living biomass from a simple scoop of dirt, map out massive evolutionary trees without needing supercomputers, and analyze complex microbiomes all simultaneously.
19:10A lot. What is the ultimate driver pushing this into the mainstream? The paper makes it abundantly clear. It all comes down to the price tag. Cost and throughput are the ultimate arbiters of accessibility in modern science.
19:21If high resolution biodiversity monitoring is going to be implemented continuously, locally, and globally to actually support the targets of the global biodiversity framework, It has to be economically feasible.
19:32For developing nations, not just massive, well-funded universities. Exactly. And the corporate technological race happening right now in short read sequencing is staggering to witness. It feels like an absolute arms race right now in the sequencing world.
19:47You read through the specs detailed in this paper, and it is a fierce battle for the future of global conservation tech. The reigning champion setting the baseline is alumina. Right. Their Novasec X platform is an absolute powerhouse, pumping out 8 terabases of sequencing data in under 2 days.
20:06They are driving the cost down to roughly $2 per gigabase of data. But the scrappy innovators are coming at them from completely different angles. Oldman Genomics, for instance, has introduced a platform called the UG100.
20:18They are tackling the cost issue by changing the fundamental chemistry. Okay, so. They utilize a mostly natural sequencing by synthesis approach. By using mostly cheap, unlabeled natural nucleotides and only a very small fraction of expensive fluorescently labeled ones, they drastically reduce the cost of the consumable reagents.
20:36Pushing the price down to around $one per gigabase. And then you have element biosciences using what they call avidity sequencing. They've essentially decoupled the physical enzymatic edition of the DNA bases from the fluorescent detection step.
20:50And this separation gives them extremely high accuracy while keeping costs highly competitive around $2 per gigabase. Notably on smaller benchtop machines that don't require massive lab infrastructure.
21:02And looking toward the immediate horizon, the paper highlights Roch's new platform called Sequencing By Expansion, or SBX. This one is wild. It involves truly mind bending molecular engineering. Instead of trying to read the microscopic DNA directly as it exists naturally, they literally expand the native DNA into much larger surrogate polymers called expandamers.
21:23They chemically stretch the DNA out to make it physically bigger and easier for the machine to read. It's like turning a tiny font into large print. Perfect analogy. These expandamers are structurally about 50 times longer than the original native DNI strand.
21:37The machine then pulls these physically expanded molecules through a highly sensitive CMOs-based nanopore array. Because the molecule is physically larger, the detection sensitivity is vastly enhanced.
21:49The projected throughput for this system is massive, and the paper notes, it could potentially drop the cost of sequencing to an astonishing ¢50 per gigabase. 50 Cent per gigabase. That is the true democratization of science and action.
22:03It really is. It means that an environmental protection agency in the developing nation doesn't need to secure a multimillion dollar international grant just to monitor their own local ecosystems. They can deploy cheap, rapid short-rid sequencing locally to track invasive species, monitor water quality, and protect their native biodiversity in real time.
22:24Furthermore, these plummeting costs directly enable massive global initiatives like the Earth Biogenome project. Which harbors the wildly ambitious goal of sequencing the genomes of all eukaryotic life on Earth.
22:36By utilizing these affordable shortread methods to fill in the massive taxonomic gaps from museum specimens and bulk environmental samples, scientists are constructing a much denser, more accurate phylogenomic framework of life.
22:49This high volume data perfectly complements the slower, more expensive, high quality, long read reference genomes. So, to synthesize this entire journey for you listening today, short read sequencing is absolutely not yesterday's news.
23:04Far from it. It might not give you the pristine, uninterrupted chapters of the DNA encyclopedia right out of the box, but it is undeniably the versatile, rugged, multi-tool of modern genetics. It thrives on the chaos of the natural world.
23:17It is the precise technology that is allowing us to map ancient dusty museum archives, decode the complex genetic soup of our oceans, and build the definitive tree of life faster, cheaper, and more comprehensively than ever before.
23:32Building these vast, highly accurate genomic databases is far more than just an academic exercise in cataloging. It is the fundamental non-negotiable prerequisite for modern conservation. It really is. is the only viable way we can effectively track, monitor, and aggressively protect our fragile ecosystems against the accelerating pressures of climate change and habitat loss.
23:52If we do not know precisely what is out there, we cannot possibly save it. And on that note, I want to leave you with one final brand new thought to ponder as you go about the rest of your day. We've talked extensively about using this technology to pull genomes out of 100-year-old museum drawers or a scoop of forest dirt.
24:11But if short red sequencing thrives on fragmented, degraded, highly damaged DNA, What happens when we start pointing this incredibly cheap, powerful technology at the ancient ice cores melting out of the permafrost right now?
24:27Oh, wow. We might not just be cataloguing the present ecosystems. We might be right on the verge of waking up the genomic ghosts of the last Ice Age, reading the shredded DNA of ecosystems that haven't seen the sun in 50,000 years.
24:38It completely changes how you view the hitter layers of the world around us. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
24:51If you enjoyed this, follow, or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
25:06Thanks for listening and join us next time as we explore more science base by base.