This paper reviews polygenic risk scores (PRS) and social determinants of health (SDoH) and outlines best practices for integrating PRS and SDoH across diverse populations to improve prediction and equity.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. Imagine taking a cutting edge DNA test to predict your risk for diabetes.
0:11Now, what if the accuracy of that test completely changed just because of the ZIP code you live in, the air you breathe or the social stresses you face daily? It sounds like science fiction, but it is actually one of the most pressing hurdles in modern precision medicine right now.
0:26Exactly. I mean, how could the conditions of our neighborhoods? literally alter the mathematical performance of our genetic blueprints. What really happens when the clean, hard data of our DNA collides with the messy, unequal reality of human society?
0:41It gets complicated very quickly. It really does. So today our mission for this deep dive is to explore the fascinating intersection of genetics and our environment and really help you understand how science is trying to combine our biology with our social realities to predict disease.
0:55And to do that. Today, we celebrate the work of Sarah J. Cromer and her colleagues at the Primed Consortions SDOH working group, who have advanced our understanding of how to integrate polygenic RISA scores and social determinants of health across diverse populations.
1:11Yeah, their work is foundational here. It really is. The paper was published in the American Journal of Human Genetics on March 10th, 2026, and it essentially acts as a definitive guide for the scientific community on navigating this incredibly complex intersection of genetics and society.
1:27Okay, let's unpack this because we are basically operating at the intersection of 2 massive predictive frameworks. On the genetic side, we have polygenic risk scores or PRS. And we all know we've moved way beyond the era of monogenic traits.
1:41We're just, you know, a single mutation dictates your outcome. Most common diseases we deal with, like type 2 diabetes or cardiovascular disease. They are highly polygenic. Which means they are influenced by 1000s or even 1000000s of SMPs, single nucleotide polymorphisms spread across the entire genome.
1:59Right, and a polygenic risk score, essentially takes the summary statistics from these massive genome wide association studies and calculates a cumorative weighted sum of an individual's risk alleges. Yeah, the weight of each variant is determined by its estimated effect size on the disease.
2:17It's this brilliant mathematical quantification of your inherent biological predisposition. But then you have the other half of the matrix. The environment, captured under the umbrella of social determinants of health or SDOH.
2:31SDOH. Exactly. These are the conditions in which people are born, grow, work, live, worship, and age. It is the extrinsic factors that dictate a person's biological exposure over their whole lifetime. So we're talking about everything from individual educational attainment to neighborhood walkability.
2:48Air pollution, systemic, social support structures, you name it. These determinants act as the underlying drivers for the vast majority of physical and psychosocial exposures a human being encounters. But integrating those 2 frameworks exposes a really severe historical and clinical flaw, doesn't it?
3:03A massive one. Because the overwhelming majority of contemporary polygenic risk scores have been built almost exclusively using data from people of European descent. Yes. The models were trained on a highly specific genetic architecture.
3:17So when researchers take a PRS optimized for European descent populations and apply it to African ancestry populations, for example. The accuracy just drops off a cliff. It drops to roughly 20 to 40% of its original efficacy.
3:3020 to 40%. That is an astonishing degradation and algorithmic performance. And frankly, it's a massive gap that could worsen health disparities if we rely on these scores for precision medicine. Absolutely.
3:43And that degradation is fundamentally driven by population genetics. It comes down to allegal frequencies and crucially, patterns of linkage to equilibrium. Right. linkage to equilibrium. Yes, the non-random association of alleles at different low si.
3:57Because modern humans originated in Africa and subsets migrated outward, African genomes possess significantly greater genetic diversity and much shorter blocks of linkage to equilibrium compared to European genomes.
4:09So when apologenic risk core, built on European data, uses specific tag SMPs as proxies for the causal variants. Those tag S&Ps often no longer correlate with the causal variants in a population of African descent.
4:24Because the blocks are shorter. Precisely. The model is essentially looking for biological signals that are no longer mathematically linked in that specific genome. And if you rely on those degraded models in a clinical setting, you risk actively harming patients.
4:39You might deny non-European patients preventative care because the algorithm miscalculated their risk. Which is why the primed consortium framework enforces a strict demarcation in the vocabulary we use.
4:51We have to separate socially constructed categorizations from continuous biological metrics. Meaning race and ethnicity. Right. The paper points out impartially that race and ethnicity are socially constructed categories.
5:03They're often based on skin color or cultural traditions, but they have no intrinsic basis in biology. They shift over time in geography. Exactly. They are dynamic proxies for a person's lived experience and where they fall within a societal hierarchy.
5:17They capture the downstream effects of historical and structural forces, which manifest as disparities in SDOH. But historically, science treated them as biological. Which was typological thinking. Putting humans into rigid biological boxes.
5:31And that completely fails to account for admixed populations. It obscures how fluid genetic inheritance actually is. So science is moving toward genetic similarity. It's a continuous spectrum of shared DNA rather than discrete continental ancestry buckets.
5:46It allows researchers to plot individuals along a multidimensional spectrum. Which solves the biological categorization problem. But it brings us to the core methodology of this whole deep dive. How do you translate the messy lived experience of SDUH into standardized computable data?
6:02That is the big challenge. Researchers measure SDOH at 2 distinct levels, the individual level and the area level. Individual level being things like direct surveys. Yes. Questionnaires, electronic health records.
6:14It isolates the proximate realities of a single person. Their specific household income, their years of education, personal insurance status. Or self-reported psychosocial stress. Exactly. But then you have area level measures, which operate on an entirely different statistical plane.
6:31They use large scale administrative databases. Like the census or the American community survey? Right. To quantify the broader contextual forces within a ZIP code or a census tract. It calculates the baseline environmental pressure of the ecosystem that individual inhabits.
6:47And the paper breaks this environmental pressure down into 4 core domains, right? Let's start with the 1st one, socioeconomic. At the individual level, this is education, employment, wealth accumulation, food and housing security.
6:59But capturing that accurately is notoriously difficult. Oh, absolutely. Income fluctuates wildly over a lifespan. And it doesn't consistently correlate with accumulated wealth, especially outside the traditional workforce.
7:10Furthermore, individual socioeconomic data suffers from severe missingness. The paper gave a really striking example of this with the all of us research program. Yes. Income data was absent for approximately 20% of their survey respondents.
7:24And it wasn't randomly missing. It heavily correlated with specific demographic variables and other adverse SDOH factors. If you try to impute that missing data, you just introduce massive analytical noise.
7:36Which is exactly why researchers are pushed toward area level socioeconomic measures. They use composite deprivation indices. Like the Townsend deprivation index. Or the area deprivation index. These tools aggregate dozens of census variables, poverty line proportions, median property values, unemployment rates, into a single standardized score representing neighborhood disadvantage.
7:58Okay, so that's the socioeconomic domain. The 2nd domain is sociocultural. This attempts to quantify the social fabric and relational stressors. Individually, researchers use validated instruments like the Berkman Symes social network index.
8:11What exactly does that measure? Social group cohesion, marital status, the breadth of your support network. The domain also measures perceived discrimination, caregiving burdens, and institutionalization, like history of incarceration.
8:24But how do you translate socio cultural factors to the area level? You can't exactly survey a ZIP code's feelings? No, but you can utilize spatial proxies for systemic discrimination. Researchers analyze the geographic concentration of specific demographic groups to measure segregation.
8:42They look at the prevalence of single parent households or the density of ethnic enclaves. Using geography as a macroscopic indicator of structural inequality. Precisely. Then we move to the 3rd domain, the physical environment.
8:55This one feels a bit more tangible. It is heavily area level dominated. Geospatial mapping, environmental monitoring. We were talking about ambient air pollution, specifically fine particulate matter, like PM 2.5.
9:07which we know drives systemic inflammation. And cardiovascular risk. They also use satellite imagery to calculate the normalized difference vegetation index, or NDVI, to objectively quantify neighborhood green space.
9:19They look at extreme heat islands, right? And the food environment. Like, how close are you to a grocery store with fresh produce versus a block packed with fast food outlets? And we have to remember, the physical environment is profoundly shaped by historical policy.
9:34Exposure to industrial zoning or urban heat islands is frequently concentrated in neighborhoods previously subjected to structural redlining. The legacy of discriminatory zoning is literally recorded in the physical hazards databases today.
9:47Exactly. Now, the final domain is healthcare access. So, insurance tiers, out of pocket costs at the individual level. Personal health literacy, yes. At the area level, it's evaluating the structural capacity of the regional healthcare grid.
10:02Using things like the area health resource file. Yes, to calculate the per capita density of primary care providers or the median public transit time required to reach a clinic. Okay, so we have PRS mapped out through continuous genetic similarity.
10:15And we have these 4 comprehensive domains of SDOH. But if I'm a scientist, my next massive hurdle is harmonization. Harmonization is incredibly tricky. It's the statistical process of reconciling data from diverse cohorts to allow for unified analysis.
10:31Because if one study measures income in $199 in rural America, and another measures income categories in 2020 urban America, how do scientists combine that data? Statistically speaking, it's treacherous.
10:45An income bracket from 1995 carries completely different purchasing power than that same bracket in 2024. If you don't account for inflation, regional cost of living and time trends, you introduce critical mispecifications into your model.
11:00There's no universal gold standard for measuring SDOH. No. Which forces researchers to build highly customized theoretical frameworks just to ensure they're comparing equivalent environmental exposures before they even touch the genetic data.
11:13But once they do harmonize it, they move to the analytical modeling. And this is where we get into effect estimation. Right. Deciphering the mathematical relationship between the genetics, the environment, and the disease.
11:23We look at 3 main analytical frameworks. Let's start with main effect estimation. Main effect estimation seeks to isolate the independent contribution of a specific exposure on the outcome. The cardinal rule here is temporality.
11:35Meaning timing matters. Because your germlin genetics are fixed at conception. So an environmental exposure later in life, like getting a high stress job at age 30, cannot biologically alter your structural DNA sequence.
11:49But the statistical models still have to adjust for confounding variables. Like population stratification, yes. And selection bias, where the SDH factors being studied might influence a person's likelihood of even surviving to participate in the cohort.
12:04Which skews the data. Okay, then we have effect modification, which is testing interactions. Yes, does the magnitude or direction of an exposures effect change depending on a 3rd variable? Like, is the penetrance of a high polygenic risk score amplified by a patient's low socioeconomic status?
12:19And the paper points out, you have to test this on two different mathematical scales. Multiplicative versus additive. This is a crucial distinction. The multiplicative scale evaluates relative risk. It asks if the combined effect multiplies the baseline probability of the disease.
12:34Generating odds ratios. Exactly. The additive scale evaluates absolute risk difference. The raw number of additional disease cases attributable to the interaction. So you can see no interaction on the multiplicative scale, but a massive interaction on the additive scale.
12:48Precisely. If an adverse SDOH exposure significantly raises the baseline risk in a vulnerable group, a constant relative genetic risk will result in a vastly larger absolute number of sick individuals.
13:00Okay, and the 3rd framework is mediation. Mediation asks if a variable acts as an intermediate step in the causal pathway. Does body mass index mediate the effect between SDOH and diabetes? It calculates the proportion of disease risk passing from, say, a poorer food environment index through the mediating variable of BMI resulting in type 2 diabetes.
13:23Which brings us to the perfect concrete example. What's fascinating here is how type 2 diabetes perfectly illustrates this entire genetic and environmental collision. T2D is highly heritable. Very complex polygenic architecture.
13:36Yet its clinical onset is overwhelmingly dictated by environmental risk factors. It is a profoundly heterogeneous disease. It manifests through entirely different physiological mechanisms depending on the population.
13:48Yes. You have pathways driving peripheral insulin resistance, often linked to adiposity, and pathways driving pancreatic beta cell dysfunction. And we see stark population variances here. Black and Hispanic individuals in the U.S. often experience T2D at younger ages and at lower BMIs compared to European ancestry populations.
14:09And East and South Asian populations frequently develop the disease at even lower BMIs, driven by distinct propensities for visceral fat accumulation and vulnerabilities in beta cell secretion. Yet the dominant PRS models for T2D were built on European cohorts.
14:25So when those scores fail in diverse populations, the paper explains it happens via two distinct paths. Path one is the differences in genetic architecture we discussed earlier. The allegal frequencies and linkage to equilibrium just don't match up.
14:39Right. But path 2 is entirely different. In path two, the genetic architecture might actually be highly similar between the populations, but the SDOH distributions are radically different. A favorable SDOH in one population might associate with completely different sociocultural factors in another.
14:55Think about poverty. In a rural cohort, poverty might cluster with lower air pollution, but severe physical isolation from clinics. While in a dense urban cohort, poverty clusters with extreme industrial pollution, high crime rates, and elevated psychosocial stress, despite being physically close to a hospital.
15:14The core socioeconomic label poverty is identical. But the compounding environmental exposures are completely divergent. that alters the outcome. The environmental variables can multiply the disease prevalence so intensely that the genetic risk core just gets lost in the noise.
15:30It becomes clinically irrelevant. Which fuels the issue of selection bias and informative missingness. The populations enduring the most severe SDOH burdens are the least likely to be represented in research.
15:41Let's talk about the all of us data again. The paper highlighted a wild statistic. In that program, over 90% of Asian responders have some college education. And 40% have advanced degrees. That is massively skewed compared to the general Asian identifying population in the US.
15:57So if you use that data set to analyze the protective effects of education on genetic risk within that demographic, your algorithm will be entirely distorted. You are observing a highly restricted, affluent subset.
16:10The data from marginalized segments isn't randomly missing. It's missing because socioeconomic barriers prevented their participation. Digital divides, historical mistrust of institutions, lack of physical access to study centers.
16:24And the disease outcome itself suffers from misclassification. Marginalized groups often face delayed T2D diagnoses due to reduced healthcare access, or they present with atypical FIA types simply because the standard criteria were optimized for European populations.
16:41So your reference panels are skewed, your environmental data is missing non-randomly and your clinical outcomes are misclassified. As a perfect storm. And if you try to fix it by just throwing geographic SDOH data into the model, you hit another wall.
16:53The modifiable aerial unit problem. The extreme sensitivity of spatial data to geographic scale. This blew my mind. Sensus track data is highly consistent, but if you aggregate that same data to the ZIP code or county level.
17:07The internal heterogeneity dilutes the variants. A single county can have heavily resourced, affluent neighborhoods just miles away from communities suffering from severe industrial pollution. If you average those extremes together, you get a mathematically moderate score that completely erases the lived reality of both populations.
17:26Applying that smoothed out area metric to an individual will fundamentally miscalculate their risk. The spatial resolution of your data has to match the scale of the environmental mechanism. So what does this all mean for the future of medicine?
17:39We are aggressively moving toward models that predict disease by combining PRS and SDOH, but we have to be incredibly careful. The paper warns specifically about using racial designations and clinical algorithms.
17:51Yes. Impartially speaking, for decades, calculators for kidney function or predicting success of a vaginal birth after caesarian used race as an independent biological variable. They applied flat mathematical race corrections.
18:03Which systematically altered risk profiles based on a socially constructed category. And applying those corrections unthinkingly obscures disparities, and often restricts minoritized populations from access and care.
18:16But the paper also warns that just ignoring social context can miscalibrate risk models too. Right. Just deleting the racial category to make a colorblind algorithm doesn't solve the statistical problem.
18:27If you remove the proxy without replacing it with accurate continuous measures of the specific structural determinants like stress or access barriers, the model will still fail vulnerable populations. And in the rush to use environmental data, researchers have to watch out for the ecological fallacy.
18:45A foundational trap from 1950. It's when you observe a correlation at the population level and incorrectly assume it applies to every single individual within that population. You cannot assume a person has a low income just because they live in a low income ZIP code.
18:58Exactly. A patient in a deprived census tract might have substantial personal wealth and excellent healthcare. Assigning them an adverse individual risk profile based solely on their geographic coordinates is a severe algorithmic error.
19:12So to fix this, the study points to the absolute need for massive non-European genomic resources. We need to diversify global biobanks to fix these predictive accuracy gaps. And researchers have to use causal language very carefully.
19:26Since SDOH metrics are indirect proxies for complex phenomena, you can't carelessly claim a specific factor directly causes a genetic interaction. Because that could unintentionally reinforce racial essentialism or stigmatize minority groups.
19:40Exactly. It requires what the authors call humble engagement with affected communities. Understanding their concerns about algorithmic surveillance and ensuring these tools dismantle disparities rather than encode them into precision medicine.
19:52It really requires us to distill this massive complexity down to its essence. To truly unlock precision medicine, we cannot look at DNA in a vacuum. We must integrate the vast overlapping complexities of our genomes with the social and environmental realities of where we live, work, and age.
20:11Isolating the genome from the ecosystem is fundamentally misunderstanding human development. So I leave you with this. What does this mean for the future of your own healthcare, when your doctor might need to look at both your sequence genome and your neighborhood's history, just to map out your medical future?
20:28It is a whole new frontier. This episode was based on an open access article under the CCBY4 license. You can find a direct link to the paper and the license in our episode description. If you enjoy this, follow or subscribe in your podcast app and leave a 5 star rating.
20:44If you'd like to support our work, use the donation link in the description. Now stay with us for an original track created, especially for this episode, and inspired by the article you've just heard about.
20:53Thanks for listening and join us next time as we explore more science based by base.