This episode reviews a GWAS of speech rhythm (prosody) perception using the TOPsy task (n≈1,501 European-ancestry), reporting 14 suggestive loci, nominal enrichment for songbird vocal-learning gene sets, and polygenic links to reading and musical rhythm.
0:00Welcome to Base by Base, the papercast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. You know, we spend, well, we spend a lot of time thinking about the actual words we use, right?
0:13The definitions of the vocabulary. the literal meaning of things. Exactly. But there's this whole other layer of communication that we process constantly. And it's totally under the radar. So consider a really simple sentence, like, John bought the car.
0:28Okay. Now, listen to how the meaning completely changes just by shifting the stress. Like if I say, John bought the car, I'm clarifying what he bought, right? Like, not the motorcycle, the car. Right, yeah.
0:38But if I say, John bought the car, I mean, now I'm clarifying who bought it. Not Fred John. And the really fascinating thing there is that the words are identical. The syntax hasn't changed at all. What shifted was the intent.
0:51And that subtle manipulation of stress, you know, the shifting of intonation and rhythm, that is exactly what we call pro city. So what really happens in our biology. When we process these subtle rhythmic cues in speech.
1:05Because today, we're asking you to consider how this seemingly simple auditory skill might actually be, well, hardwired into our genetics. It's incredible to think about. It really is, how it's inextricably linked to our ability to read, and remarkably how it connects to the evolutionary history of songbirds.
1:26Songbirds. I mean that's the part that really gets me. I know, we'll definitely get to that. But I like to think of speech rhythm as sort of like the invisible structural beams of a house. You don't actively look at the beams when you walk into a room.
1:38No, you're looking at the furniture or the paint. Exactly. But without those beams, the whole structure of communication just simply collapses. Yeah, and if those beams are even slightly off kilter, the house might still stand, but the experience of living in it is different, or in this case, the experience of communicating and learning is just fundamentally altered.
1:56So today, we are pulling the blueprints for those structural beams. We're doing a deep dive into the underlying biological architecture of how we actually hear the rhythm of speech. And today, we celebrate the work of Alyssa C. Scartotsi, Shrishti Nayak, Reina L. Gordon, and their extensive collaborative team at Vanderbilt University Medical Center, and other institutions, who have advanced our understanding of the genetic architecture of pro city perception.
2:21It really is an incredible genome wide investigation. It was published in the journal HGG Advances in 2026. It's a huge step forward. It is. So let's unpack the core concept first. We define pro city as the intonation and durational patterns in spoken language.
2:37Right, the melody of speech basically. Yeah. But this team narrows in on a very specific slice of that, which they call speech rhythm perception. Right. So speech rhythm perception focuses strictly on recognizing the stress patterns you just demonstrated with the car example.
2:51It's the ability to detect variations in emphasis within words or sentences and, you know, larger spoken structures, and the sensitivity actually emerges when we are infants. Wow, really, so long before we even understand vocabulary.
3:06Exactly. Infants are tracking these rhythms, but it becomes hypercritical during mid-childhood and elementary education. Because, well, that is the exact window when we were learning to read. Yes. Children rely heavily on these speech rhythm cues for crucial linguistic processes.
3:23So take lexical retrieval, for instance. Which is what exactly? When you hear a word, your brain has to quickly fetch the correct meaning from your mental dictionary. Prosody helps narrow down that search.
3:36Oh, I see. It gives you a hint based on the sound. Right. And it also aids in morphosyntactic parsing. Okay, that's a lot of syllables. I know, it's a highly technical way of saying, your brain decodes the grammar of a sentence on the fly.
3:49You're using the rhythm to figure out where the commas and periods belong, even in spoken conversation. That makes so much sense. I want to linger on something from the source material, though, because it completely changed how I think about reading.
4:01Oh, the implicit prosody hypothesis. Yes. The idea that speech rhythm perception is not just about spoken language. Like, think about reading a text message from a friend that is typed in all C cap BS.
4:15Oh, yeah, you definitely hear it differently. Versus one typed in all lowercase, right? Your brain completely changes the volume, the rhythm, and the pitch of the voice in your head. It's practically shouting at you Exactly.
4:27Even if you are sitting alone in a dead silent room reading a book silently to yourself, your brain is actively projecting prosodic changes onto the text. You hear the emphasis, the pauses, the rhythm internally.
4:40Because you are generating that structural framework inside your own mind to make sense of the symbols on the page. And that internal acoustic simulation is exactly why speech rhythm perception skills are highly predictive of a child's success in learning to read.
4:54Because they have to build that internal voice. Exactly. If you struggle to hear the rhythm and spoken words, you will likely struggle to project that necessary rhythm onto silent text. And this correlation, it actually persists all the way into adult literacy.
5:07Which brings up the genetics, because Quinn and family-based studies have shown us that speech and language traits generally have moderate to strong genetic influences. They're very heritable, yeah. Yeah, we were talking about heritability estimates ranging from .46 to .97 for related traits like phonological awareness.
5:27Which is huge. It tells us that a massive portion of our language ability is inherited directly from our parents. But wait, I am, I'm struggling with this a bit. Well, if this trait is so critical to human communication, and if it literally dictates how well we learn to read.
5:44And if we already know language traits are highly heritable, why on earth did it take us until 2026 to map the genetics of prosody? It's a great question. Shouldn't this have been like one of the 1st things genetics has found when we mapped the human genome?
5:59You would think so. But the delay really comes down to a fundamental bottleneck in scientific measurement. A bottleneck in how we test people. Yeah. Historically, measuring a complex behavior like prosody perception in a massive population has been nearly impossible.
6:14When you want to find genetic signals for a behavioral trait, you need 1000s, often 100s of 1000s of people. Right, for statistical power. Exactly. If you are studying cholesterol, it's easy. You take a blood sample from 100,000 people.
6:28But you cannot pull 100,000 people into a lab one by one, sit them in a soundproof booth, and have a trained researcher manually score how well they hear syllable stress. No, the logistics are just impossible.
6:43It would take decades. And that is the frustrating trade-off and genetic research. You can have incredibly rich, deep phenotyping. Meaning you measure the trait perfectly in a highly controlled lab. Yes, but only for a tiny group of people, or on the flip side, you could have a massive Samuel size of 100s of 1000s of people, but your measurement of the trait has to be shallow, like a multiple choice self-reported survey.
7:07Which doesn't really capture something as nuanced as hearing rhythm. Exactly. Because of that bottleneck, the specific biology of prosody has just remained a black box. Researchers simply couldn't get a pure measurement at scale.
7:20Well, this brings us to the core methodology of this deep dive, and how this specific team finally broke open the black box. They really do. They designed a genome wide association study, or GWS, involving 5,501 individuals.
7:37These were individuals of European genetic ancestry, and to solve that historical measurement bottleneck, they utilized an innovative piece of technology called Topsy. Yeah, Topsy. It stands for the test of prosody via syllable emphasis.
7:50I love that acronym It's great. And it completely bypasses that traditional trade-off between rich phenotyping and large sample sizes. Because it's fully remote. Exactly. It's a 28 item remote auto scored test.
8:02Participants just log on from anywhere and listen to multisyllabic spoken words. And then they are simply asked to identify which syllable carries the primary emphasis. But the genius of Topsy isn't just that it is remote.
8:14It's actually the engineering of the test itself. Right, how they design the audio. Yeah, the sources note that teopsy specifically counterbalances syllable length and syllable stress position across all the items.
8:25Which is absolutely vital. If one syllable is physically longer than the other, a person might identify it just because it took longer to hear. Oh, I see. Not because they actually perceived the linguistic stress.
8:36Exactly. They might just be reacting to the duration. So by balancing those factors, the test isolates the exact cognitive skill being measured. It is basically like a set of noise canceling headphones for geneticists.
8:49I love that analogy It blocks out all the background static, you know, of a person's general intelligence or their vocabulary size or simple auditory reaction time. It allows the researchers to capture only the pure isolated frequency of rhythm perception.
9:03And with that pure measurement finally in hand, the researchers were able to map those teopsy scores against approximately 7000000 genetic variants across the participants' genomes. 7 million. That's a lot of data.
9:16It is, but looking at the human genome in isolation really only tells part of the story. So they took a brilliant leap here. They employed a process called gene set enrichment analysis. Okay, before we throw out too much jargon, let's break down what they actually did here, because this is where this deep dive gets wild.
9:34It really does. They essentially took the genetic profile of our human listeners and overlaid it onto the brain scans of singing birds. Just to look for matches, yeah. They looked specifically at the zebra fin.
9:46Zebrafage, okay. They wanted to know if the human genes associated with perceiving speech rhythm actually overlap with the genes that are differentially expressed in the brains of zebra finches while they are actively singing their complex songs.
10:00That is just mind blowing. And on top of comparing humans to songbirds, they also use polygenic scores, or PGS, to compare our prosody genes against other human traits. Right. So think of a polygenic score as a mathematical summary of your genetic predisposition for a specific trait.
10:17Like a single number that represents your genetic risk or ability. Exactly. And because they had this brand new genetic snapshot for prosody, they could check it against massive existing genetic databases.
10:30So what do they compare it to? Well, they pulled data for word reading, which had a sample size of over 27,000 people. Okay. They pulled data for a voice pitch variability. And they pulled data for beat synchronization, which is the ability to clap along to a musical beat.
10:46And how big was that data set? That had data from a staggering 600,000 people. Wow. Okay, let me see if I have the full picture here. We have a noise canceling measurement tool in Topsy. We have a cross species comparison with zebra finches.
11:00And we have mathematical modeling against huge data sets of human reading and musical rhythm. That's the setup, yeah. So what did the data actually reveal when they put it all together? Well, starting with the genome wide association study on those 5th 10501 individuals, they did not find any single genetic variant that surpassed strict genome wide significance.
11:19Which, to be fair, makes complete sense in the genomics world. Right. For a complex behavioral trait. A sample size of 1500 is actually relatively small. Yeah. You usually need 10s of 1000s of participants to cross that extremely strict statistical threshold and say, this one specific genetic letter definitively causes this trait.
11:39Exactly. The study was underpowered to find undeniable single variant signals. However, they did identify 14 suggestive significant signals. The suggestive meaning why. These are distinct genetic low side that show a very strong statistical trend, meaning they are highly promising leads, even if they haven't crossed that final threshold.
11:59Got it. And when the team mapped these 14 signals to functional genes, some absolutely fascinating biological mechanisms emerged, particularly around 2 genes, TTLO1 and GP2. Let's dive into those mechanisms.
12:11How does a gene like TTLL1 actually influence our ability to hear the stress in the word car? So TTLO1 is intimately involved in the organization of microtubules. Okay, microtubules. To understand why that matters, picture a single neuron in your brain as a long, sprawling highway.
12:26Okay, I'm picturing a highway. Microtubules are the physical scaffolding and the internal transport tracks of that highway. They maintain the structural integrity of the cell. So they keep the roads paved and clear.
12:38Right. And if those microtubules are compromised, the neural signal traveling down the highway might be delayed, just by a few milliseconds. And when we are talking about perceiving the rhythm of speech, milliseconds are literally everything.
12:52Everything, rhythm is pure timing. Right. If the brain's internal highway system has structural delays, the ability to process rapid rhythmic auditory cues is just going to degrade. Precisely, which is why seeing a neurodevelopmental gene like TTLL1 pop-up is so compelling.
13:09It links the behavioral trait of listening directly to fundamental brainwiring. Wow. And then we have GP2. The sources note that this one has previously been associated with endocrine traits and sleep quality.
13:19Sleep quality. Yeah, which, at 1st glance, feels completely disconnected from how we process language. I mean, sleep and speech. They seem unrelated. But if we think about the mechanisms. Are we just looking at different scales of biological timing?
13:32That is a brilliant connection to make. Think about it. A circadian rhythm. Our sleep wake cycle is basically a biological oscillator operating on a 24 hour loop. A heartbeat is an oscillator operating on a 12nd loop.
13:46And speech rhythm perception involves processing oscillators at the millisecond level. Oh wow. And the genetic architecture for clapping to a musical beat is also known to correlate with traits like circadian chronotype in insomnia.
14:00So GP 2 suggests there might be a deep underlying biological rhythmicity. Like, the same genetic clockmakers in our body might be responsible for telling us when to go to sleep, when to clap our hands to a song and how to process the rhythmic stress of a spoken sentence.
14:14Yes. The biological overlap is profound, and honestly, that overlap becomes even more astounding when we return to the zebra finches. Right, the bird brains. Exactly. The gene set in Richmond analysis revealed a highly significant intersection.
14:27The human genes associated with speech rhythm perception were significantly enriched in a specific region of the songbird brain called Area X, and specifically during their singing behaviors. Let's explain Area X for the listeners.
14:40Why is that specific region, the smoking gun here? So, Area X and birds is the known evolutionary ortholog of the human basal ganglia. In ortholog, meaning it evolved from the same ancestral structure, right?
14:55Exactly. In humans, the basal ganglia is a deep brain structure, famous for motor control. Like it is what gets compromised in Parkinson's disease. Right. We hear about it a lot with movement. But it is also a massive hub for processing musical rhythm and precise neural timing.
15:10It helps your brain anticipate the next beat in a song. Or anticipate the end of a sentence based on the stress of the words. Yes. So the exact same genetic machinery, firing in the basal ganglia of a human who is trying to figure out if John bought the car is firing an area X of a finch when it is learning to sing its complex song.
15:28That is a stunning example of convergent evolution. really is. But the findings from the polygenic score analysis bring this cross species biology right back to our everyday human lives. The PGS results were amazing.
15:40They were. The data showed that your genetic predispositions for word reading, for beat synchronization and for voice pitch variability. All of those actively and positively predict your score on the Topsy speech rhythm test.
15:56They map right onto each other. I want to make sure the listener is truly internalizing this math. If we had your genetic data right now. The specific genes that make you naturally good at clapping along to a song on the radio can mathematically predict your ability to hear the subtle linguistic stress in a spoken word.
16:13It's wild. The data confirms a concept known as the may PLE framework that stands for musical abilities, pleotropy, language and environment. Pleotropy being the key word there. Right. Pre entropy is when one single gene moonlights, doing multiple seemingly unrelated jobs across the body.
16:29This study provides concrete biological evidence that rhythm processing is a domain general trait. Meaning rhythm is rhythm. The brain doesn't care if it is a drum solo or a conversation. It uses overlapping genetic and neural architecture to process the timing of both.
16:43And incredibly, it uses that exact same architecture to help us decode written text on a page. Which brings it all full circle. When we step back and look at the whole picture, You know, the suggestive genes dictating microtubial highways, the finches using their basal ganglia, the mathematical link to collapping and reading it provides profound support for the revised vocal learning hypothesis.
17:07The revised vocal learning hypothesis. If this rhythm processing is so deeply embedded in our biology, it really makes me wonder how it all started. Right, like where did language come from? Yeah, did early humans just suddenly start using complex grammar?
17:21Well, the revised vocal learning hypothesis, posits that are incredibly complex human language with all its intricate syntax and grammar, did not just appear out of nowhere. That makes sense. Nature usually builds in steps.
17:33Right. Humans and songbirds are both vocal learners. We modify our vocalizations based on our environment. The theory suggests that language evolved convergently from early, pre-adaptive, prosody rich proto language.
17:45So like early hominids using rhythmic grunts or pitch changes and melodic vocalizations to communicate emotion or mating intent or danger long before they had actual dictionary words. Exactly. Those prosodic musical cues facilitated group survival.
18:01Because hearing and mimicking those rhythms was quite literally a matter of life and death. Wow. The genetic architecture for auditory vocal processing was deeply conserved and built upon over 1000000s of years, eventually serving as the biological foundation for complex language and much later, reading.
18:19So we are basically walking around with the ancient survival software of rhythm processing, and we use it to read paperback novels on the train. That is exactly what we're doing. That is amazing. But the implications of this study go far beyond evolutionary history.
18:33There is a massive clinical impact here today, isn't there? Huge. Let's consider neurological and mental health conditions. We have known for a long time that impaired prosody is a feature of several conditions.
18:43Like what kind of conditions? Well, individuals on the autism spectrum often exhibit differences in prosody. We also see altered prosody and clinical depression, following right brain damage, and even in the progression of Alzheimer's disease.
18:57And the source material highlights a really critical shift in how we handle this clinically. It's a paradigm shift really. Historically, doctors and scientists have focused on prosody production in these disorders, right?
19:08Meaning how flat or atypical a patient's voice sounds when they speak. Yes. But this study shifts the spotlight to prosody perception, how they are biologically processing the auditory world around them.
19:21And that shift from production to perception changes everything about patient care. How so? Because if we understand the genetic vulnerabilities underlying how someone perceives speech rhythm, it opens up entirely new avenues to understand their sensory experience.
19:35Take dyslexia, for example. Okay. Individuals with dyslexia often struggle significantly with manipulating word stress and precise auditory timing. If we know the specific genetic and cellular roots of that struggle.
19:48Like those microtubial highways we discussed earlier. Exactly. If we know that, we can design biological or behavioral interventions that target the actual mechanism, not just the symptom. The potential for early intervention, there is just huge.
20:02But as with all science, we also have to view these findings through a lens of scientific rigor. Always. As we noted earlier, the sample size here was a major limitation. We must candidly address the constraints of the study.
20:14A sample size of 4501 is fundamentally underpowered to declare absolute definitive genome wide signals for a trait this complex. Right. The 14 signals they mapped are suggestive. They are excellent leads, but they must be replicated in much larger cohorts to become undeniable fact.
20:33And there's also a glaring diversity issue in the data set. Yes. 88% of the sample used for the primary GWs consisted of individuals of European genetic ancestry. Which is a systemic ongoing problem in genomics.
20:45It is. And human language is incredibly diverse, with vast differences in how rhythm and tone are used globally. Think about tonal languages, for instance. Right. To truly understand the universal genetic architecture of human communication.
20:57We absolutely need to deploy tools like tideopsy in larger, far more diverse bio banks that actually represent the global population. Exactly. We cannot map the biology of human speech based on a single slice of genetic ancestry.
21:11But the beauty of the totopsy test is that it was designed specifically to be scalable and remote. Right, which makes that kind of diverse global deployment possible for the 1st time. So as we bring this deep dive to a close.
21:24How do we distill all of these mechanisms, the birds, the genes, and the clinical hopes into one core takeaway? Well, our ability to perceive the rhythm of speech is not just a quirky secondary listening skill.
21:35It is a deeply biological trait. It shares a common genetic architecture with how we learn to read, how we feel musical beats, and even how birds learn to sing, highlighting an incredible evolutionary convergence across 1000000s of years.
21:48It really completely changes how you experience a simple conversation, which leaves us with a final thought to mull over. What does this mean for how we might one day diagnose and support children with reading or communication disorders long before they ever open a book?
22:02This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a five-star rating.
22:17If you'd like to support our work, use the donation link in description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
22:27Thanks for listening, and join us next time as we explore more science, based by base.