In UCLA ATLAS EHR-linked biobank analyses, random forest-derived enrollment probabilities and inverse-probability weighting increased replication of known GWAS variants and altered PGS associations.
0:00So, think about the last time you went to the doctor, right? You're sitting there and they're about to draw your blood. And usually there's this little digital form on the clipboard or maybe on your patient portal app.
0:13Right, the consent form. Exactly. And it asks if they can keep your leftover blood samples for research. It seems like such a, you know, harmless, even noble thing to do. Yeah, totally. You just check yes and feel good about it.
0:25Right. You check, yes, you get your blood drawn, and you go home thinking you just helped like cure a disease or something. Which is the goal ideally. But, and this is the big question for today. If precision medicine is the future, right?
0:37If we are relying on the DNA of 1000000s of people to tailor treatments and predict diseases and train medical AI, what happens if the DNA we are starting doesn't actually represent humanity? Yeah, well, what's fascinating here is that we are discovering a profound vulnerability in modern science.
0:56I mean, we are building the future of healthcare on these massive databases called biobanks, but researchers are increasingly realizing that the invisible everyday choices of, you know, who opts in and who doesn't.
1:10It's completely distorting our understanding of human genetics. It's like, imagine trying to figure out the average American's fitness level, but you only survey people walking out of a gym. Oh, that's perfect way to put it.
1:21Right. That's essentially what's happening in genomic research right now. Yeah, it really is Welcome to Base by Base. The paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app.
1:33So glad to be here. Today we celebrate the work of the UCLA ATLA's Community Health Initiative, who have advanced our understanding of how massive genetic databases can be fundamentally biased. They really did some incredible work on this.
1:48Yeah, they did. And today, we were taking a deep dive into this really fascinating 2026 paper. They managed to actually quantify this distortion. And not just quantify it. They amazingly built a mathematical time machine to reverse it.
2:03Which is wild. Okay, so give us some background. What's the deal with these bio banks? Well, for a long time, the scientific community has been completely obsessed with what you might call the biobank gold rush.
2:14The gold rush? Yeah, basically, bio banks that link an individual static genetic sequence to their electronic health records, their EHRs. The medical files. Exactly. Those are considered the absolute holy grail for research right now.
2:29Because you don't just have a snapshot of someone's DNA. Right. You have the context. You have it mapped directly against the longitudinal story of their actual health over years or even decades. Which is incredibly powerful.
2:42I mean, instead of just asking people a survey question like, uh, do you have heart disease? Which people lie about? Or just mismember? Exactly. Instead of that, a researcher can literally see their cardiology workups, their prescriptions, their blood pressure trends over a 10 year period.
2:57Yeah, all the raw data. And then they cross-reference all of that rich real-world data with their genome. It is incredibly powerful. But, um, it comes with a massive catch. The inclusion bias. Yes, exactly.
3:09Inclusion bias. And this is really a combination of 2 things. First, you have recruitment bias, meaning who even gets asked to participate in the 1st place. Right. And second, you have participation bias, meaning out of those asked, who actually takes the time to say yes?
3:24Who clicks the button? Exactly. Now, we have known about bias in population level biobanks for a while. Like the UK biobank. Yeah, the UK Biobank is a great example. That usually manifests as a healthy participant bias.
3:39Oh, because healthy people are the ones signing up. Yeah. The people who volunteer for those large national surveys out of the blue tend to be healthier, wealthier, and frankly, they just have more free time.
3:49That makes sense. But EHR linked bio banks, like the ones attached to major hospital systems. They have the exact opposite problem. Right, because you only get asked to join a hospital's biobank if you are actually physically interacting with that hospital.
4:03Precisely. They rely on opt in consent during clinical visits. So, almost by definition. These biobanks are heavily recruiting sick people, or at least people who interact with the healthcare system constantly.
4:17Yeah. The UCLA, ATLA's community health initiative wanted to know exactly how much this specific brand of bias was skewing their genetic data. And if they could fix it? Exactly. More importantly, if they could correct it.
4:29Okay, let's unpack this because to understand the sheer scale of this problem, we really have to look at the numbers they were working with. numbers are pretty staggering. They really are. So the ATLA's initiative operates within the UCLA health system, which is this massive network.
4:44The researchers looked at a background population of over 1500000 eligible individuals who received care there between 2016 and 2023. Right, one.5 million. But out of that massive pool, only about 104,516 people actually enrolled in the ATLA's bio bank.
5:03Which is a tiny fraction. Right. So the immediate question is, who are these 104,000 people? Are they just a perfect random microcosm of the 15 million? Not even close. I mean, when the researchers profile the average participant, The differences were just mind blowing.
5:17What was the biggest difference? The single most predictive factor was simply whether a person received their primary care at UCLA. Wow. Just the primary care location. Yeah. Over 70% of the enrolled individuals received primary care there compared to just under 22% of the unenrolled population.
5:36I mean, that makes total logistical sense, right? Oh, absolutely. Because if your primary care doctor is stationed at UCLA, you are physically in their buildings more often. You're walking past the posters.
5:48Exactly. You are logging into their specific patient portal more often to check messages or test results. You are simply exposed to their recruitment materials at a much higher rate. And this naturally translates into overall healthcare utilization.
6:02Enrolled individuals, visit the doctor almost twice as much as those who don't enroll. Twice as much. Yeah. They average about 12.8 clinical visits per year. That's more than once a month. Right. Compared to 6.7 visits for the unenrolled.
6:16So, the bio bank is essentially a club for, like, frequent flyers in the healthcare system. That's exactly what it is. But the SKU wasn't just about how often they saw a doctor. Demographics too, right?
6:28Yeah, the demographic makeup was totally different. The enrolled participants were significantly more likely to self-identify as white or non-Hispanic. Okay. It was about 57% of the enrolled group versus 43% of the unenrolled.
6:42And they were also less likely to use state insurance. Wow. But, you know, the detail that really caught my eye in this deep dive was the counterintuitive finding about smoking. Oh yes, the smoking paradox.
6:54Yeah. There were significantly more previous smokers in the biobank. but fewer current smokers. It seems like a paradox at 1st glance. It really does. You'd think current smokers would be at the doctor more.
7:05But it actually ties perfectly back to healthcare utilization. Think about it. Previous smokers often have lingering health issues that require ongoing monitoring. Right, like COPD. Exactly. Early stage COPD, or they need regular cardiovascular checkups.
7:20Because they are actively managing the consequences of past smoking, they are in the clinic more often. Oh, and so they get asked to join the biobank more often. Exactly. Whereas current smokers, statistically speaking, might be avoiding the doctor altogether.
7:34Right, out of fear of bad news. Or simply prioritizing other things over preventative care. That paints a very specific picture of the type of person volunteering their DNA, but honestly, my absolute favorite detail in this whole profile is what they found hiding in the medical billing codes.
7:50Ah, yes. The ICD 10 codes. Yes. The Z-Drocumbent anomaly. It's so revealing. It really is. So every time you go to the doctor, the hospital assigns a standardized billing code to your visit, so your insurance knows what to pay for.
8:03The researchers found that enrolled individuals had a massive enrichment for 0 codes. The odds ratio was an incredibly high 7.57. Which is a huge spike. Massive. And 0 basically means encounter for general examination without complaint.
8:17It is a universal code for a routine checkup. And crucially, what happens at a routine checkup. They draw your blood. Exactly. Routine checkups almost always involve routine blood draws. Which is the exact raw material the biobank needs to function.
8:32The entire system relies on leftover blood from clinical lab tests. So if you are the kind of person who goes in for your annual physical, gets your cholesterol checked, and proactively manages your health, you are generating the residual blood required to sequence your DNA.
8:49Precisely. So, I want you listening to think about your own medical history for a second. Are you the type of person who diligently goes in for every routine checkup? Or do you wait? Yeah, or do you wait until something is visibly painfully wrong?
9:03Because your personal answer to that behavioral question might literally determine whether your specific DNA is currently shaping the future of medicine. That is the core dilemma here. The biobank is overflowing with health anxious people.
9:16highly proactive people, and people with chronic conditions requiring frequent monitoring. It's not the general public at all. It is not a reflection of the general public. So, the researchers face this monumental mathematical challenge.
9:29How do you correct a data set of 104,000 highly specific types of people so that it accurately reflects the diverse 1500000 people in the background? And their solution is essentially to build, like, an algorithmic mirror to reflect reality.
9:47They used machine learning, specifically something called a random forest classification model. Right. Right. For those who might not be familiar, a random force model works by building 100s or 1000s of digital decision trees.
9:59Like a massive flowchart. Exactly. You feed it a bunch of data, and it learns to make a series of yes or no splits to categorize things. In this case, they fed the AI, all the demographics, the healthcare utilization rates, the ZO billing codes, the insurance types, everything they had on the 1500000 people.
10:14And the model learned the exact combination of traits that lead someone to enroll. And I got really good at it, right? Incredibly accurate. It achieved an AROC score of .85. Which means what exactly? It means it had about an 85% probability of correctly distinguishing a biobank participant from a non-participant based solely on their medical record.
10:37Wow. So the AI now knows exactly what a textbook biobank participant looks like. Yes. But how do you use that knowledge to fix the genetic data they've already collected? I mean, they already have the DNA.
10:51Right. So you use a statistical technique called inverse probability waiting. Okay. Inverse probability waiting. Yeah. Once the model calculates the probability of you enrolling, they basically flip that number upside down.
11:01Okay I think we need an analogy here. Let's use one Imagine you are trying to take a poll to understand the needs of a town and you invite 100 people into a room. Okay, 100 people? 99 of them are wealthy business owners, and one of them is a minimum wage worker.
11:17If you just tally the votes straight up, the business owners completely ground out the worker. It's not a fair representation of the town at all. Right, because the worker actually represents a massive portion of the town's population that just, you know, couldn't take time off to attend the meeting.
11:32Exactly. So, to fix the poll mathematically, you downweight the business owners, you say each of their votes only counts for a fraction of a point. Oh, okay. And you upweight the single worker, you multiply their vote by 50 so that their single voice accurately represents the 50% of the town that is missing from the room.
11:52That is so smart. That is inverse probability waiting. Exactly. Yeah. If you are a wealthy white patient who visits their UCLA primary care doctor 12 times a year, your genetic data gets mathematically shrunk.
12:07Because the database is already flooded with people exactly like you. Right. But if you are someone who rarely visits the doctor, maybe you state insurance, and you just happen to randomly opt in one day after an ER visit.
12:18Your data gets upweated. Wow. Your single genetic profile is suddenly speaking for 1000s of missing people in the background population who share your profile, but never enrolled. Exactly. Your voice in the aggregate data becomes much louder to compensate for your absence in the actual enrollment numbers.
12:34Wait, hold on, though. Because I was looking at the methodology section. And something about this really tripped me up. I get the theory, but the math creates a really wild paradox. The paper states that by applying these inverse probability weights, the effective sample size of their genetic study dropped dramatically.
12:52They started with over 48,000 genotyped people who had mapped health records. But after they applied this waiting system, the effective sample size shrank down to just over 11,000. It did. That is a 4.3 fold reduction in data.
13:08Isn't the golden rule of modern statistics that more data is always better. How does throwing out that much statistical weight not completely ruin their power to discover things? Well, if we connect this to the bigger picture of data science, it reveals a really beautiful truth about statistics.
13:25Okay, lay it on me. Yes, the golden rule is usually more data is better. And shrinking the sample size by a factor of four should severely hurt your statistical power. But in this specific case, the data they're down waiting is highly redundant and fundamentally biased.
13:42By throwing out that statistical weight. They are mathematically removing the systemic noise of participation bias. And when you do that, the actual true signal of the genetics gets clear. I see. Because it's not about how loud the music is playing.
13:57It's about whether you can actually hear the melody over the static. That's a great way to put it. The overall volume is lower, that you reduce sample size, but the static of the frequent flyers is gone.
14:08allowing the quiet, true genetic associations to finally be heard. Okay, here's where it gets really interesting. Yes. So the math works in theory, and the analogies make perfect sense, but did they actually test the synthetic weighted population against real world genetics to see if it holds up?
14:25They absolutely did. They needed to prove this wasn't just, you know, a clever mathematical trick. Right. So they used a massive catalog called PGRM. What's? It lists 1000s of well-established known genetic variants, GYS variants that have already been confirmed by previous global studies.
14:40Okay, so these are like the gold standard genetic truths. Exactly. They wanted to see if their weighted, reduced size biobank could find these known associations better than the raw, biased biobank. And the results were just stunning.
14:55They really were. The weighted model, the one with the artificially shrunken sample size, replicated 54% more known genetic associations than the unawighted model. 54% more. That is a massive Lieben discovery power.
15:09It's huge. It proves that adjusting for human behavior actually uncovers biological truth. To give some concrete examples, The weighted model successfully found a specific variant in the PPRG gene that is heavily linked to type 2 diabetes.
15:23It also found variants in the CLSR2 gene cluster that are strongly linked to coronary atherosclerosis. Both of which are incredibly important genetic targets. For treating massive public health issues, and both of them were completely missed by the unweighted model.
15:38The noise of the proactive patients who get their cholesterol checked every 6 months, completely drowned out the true genetic signals for diabetes and heart disease. And that leads to a phenomenon they highlight that is perhaps the most nuanced finding in the entire paper.
15:53Okay. They refer to it as discordants, but we could think of it as the weird flip. The Weir Flip, I like that. Sometimes applying these weights didn't just uncover hidden signals. It actually flipped the direction of the genetic effect entirely.
16:07Wait, that sounds impossible. How can a gene go from increasing your risk of a disease in one model to decreasing your risk of a disease in the other, just because you multiplied the data by a weight? It comes down to how genetics interact with ancestry.
16:20And who was missing from the room? Oh. Remember how we discussed that individuals of non-European ancestry were significantly less likely to enroll in the biobank? Because of that. They received much higher weights from the algorithm to compensate for their absence.
16:37When those underrepresented individuals were upweighted, their specific ancestry driven genetic effects suddenly had a proper voice. And genes don't act in a vacuum, right? A specific genetic variant might have a totally different effect depending on the surrounding genetic background.
16:52Or environmental factors common to a specific ancestral group, precisely. This is known as linkage to equilibrium. Linkage dis equilibrium. Yeah, where genes are inherited in different blocks, depending on ancestral history.
17:06So, if a genetic variant increases risk in a European population, but decreases it in an East Asian population. And your biobank only recruits Europeans. You are going to draw a completely false conclusion about humanity as a whole.
17:19By ignoring the missing populations, the bias model wasn't just muting the truth, it was actively distorting it. Pointing science in the exact wrong direction. That is chilling. I mean, if you are a pharmaceutical company designing a precision medicine drug based on that unweighted genetic target, you might spend 1000000000s of dollars designing a drug that literally does the opposite of what you want for certain ancestral groups.
17:44It's a huge risk. And this distortion gets even more complicated when we move away from single genes and look at complex traits. Right, which brings us to the 2nd major part of their findings. Polygenic risk scores or PGS.
17:57For those who might not work with them every day, a polygenic risk score isn't about finding one smoking gun gene, like, you know, a mutation for Huntington's disease. No, it's much broader. It's about looking at thousands, sometimes 1000000s of tiny genetic variants scattered across your entire genome, and adding them all up to calculate your overall cumulative risk for a complex disease.
18:20Exactly. And the researchers ran these scores across the entire medical record. Which is called a fee wall. Right? A phenami wide association study. They specifically looked at 5 traits, major depressive disorder, BMI, rheumatoid arthritis, type one diabetes, and a behavioral trade called general happiness with health.
18:40And this is where we see the landscape of the data fundamentally shift. It really does. For traits with very strong, well-defined genetic architectures, like major depressive disorder, BMI and rheumatoid arthritis, the core association stayed strong, regardless of whether the data was weighted or not.
18:57So the underlying biology there was loud enough to cut right through the participation bias. Yes, but the behavioral traits. That's where the illusions appeared. The happiness mirage. This completely blew my mind.
19:07It's wild, isn't it? When they looked at the polygenic risk score for general happiness with health using the raw, unweighted model. They found dozens of significant genetic associations across the medical record.
19:19If you just published that paper, it would look like you had found the genetic secret to why some people are just fundamentally happy and satisfied with their health. But once they apply the algorithmic fix.
19:29Once they accounted for the fact that happy, proactive, health literate people are the ones who volunteer for biobanks in the 1st place. What happened? Almost all of those genetic associations completely vanished.
19:43Wow. They were total mirages. So they weren't genetic links to happiness at all. They were just genetic links to the types of privileged, highly engaged people who have the time, energy, and health literacy to click, I agree on a digital consent form.
19:59Exactly. The genetics of happiness were perfectly masquerading as the genetics of biobane participation. That is insane. And conversely, look at what happened with type one diabetes. Right. What happened there?
20:11In the raw biased data. The genetic associations across the medical record were weak and messy. But once the weighted model cleared out the noise of the frequent flyers, a whole cluster of new, true genetic associations for type one diabetes suddenly appeared.
20:25So the participation bias had basically been bearing them alive. Totally be them. So what does this all mean? We cannot trust biobank data at face value, right? Especially when we are looking at complex behavioral or metabolic traits.
20:40No, we really can't. The data is literally lying to us by omission. This paper serves as a massive wake-up call for the entire field of genomics. I mean, we are rapidly moving toward a future where your doctor might look at your genetic risk score to decide which blood pressure medication to prescribe.
20:57Or what age you should start getting cancer screening. Exactly. But if we build those critical clinical tools on biased data, those tools will only be accurate for the specific subset of people who opt into biobanks.
21:10Which, as we've seen, are predominantly white, heavy healthcare utilizers. Right. We risk hardwiring healthcare disparities directly into our most advanced medical algorithms. But it's not like we can just take the UCLA mathematical fix and magically apply it everywhere to solve the problem, right?
21:25No, unfortunately not. And the authors are very transparent about this limitation. This specific study and the exact weights they calculated are entirely unique to the UCLA ATLA's biobank. Because every hospital is different.
21:38Exactly. Every single hospital system has different demographics, different patient portal software, and different ways they ask for consent. Oh, that makes sense. Therefore, every biobank in the world needs to run its own complex ad hoc analysis to figure out its own unique flavor of inclusion bias.
21:55That sounds like a lot of work. It is. Furthermore, it's worth noting that while this inverse probability waiting method works beautifully for common genetic variance, it actually struggles with very rare genetic mutations.
22:07Why is that? When you shrink the effective sample size so drastically, from 48,000 to 11,000, you risk accidentally erasing a rare variant from the data set completely. Right, because if only 3 people have it and they get downweighted or filtered out, it's gone.
22:23Exactly. So it's not a silver bullet, but it is a vital, necessary step forward to clean up the data we are relying on. It absolutely is. It makes you realize just how interconnected human behavior and biological discovery really are.
22:36The next time you see a headline saying, scientists find the gene for X, you have to wonder, did they find the Jane? Or did they just find the people who like talking to scientists? That's a great point And it raises a really important dilemma that I hope everyone mulls over today.
22:51Yeah. We've established that to fix our medical data, we need these incredibly invasive AI driven models that profile our habits, right? Tracking exactly how often we go to the doctor, what insurance we use, what billing codes we generate.
23:06Just so researchers know how to mathematically weight our DNA. But if the future of precision medicine requires this level of algorithmic surveillance just to correct its own biases. Oh, wow. At what point does the cure for bias data become a massive privacy issue of its own?
23:21That is a brilliant point. Are we comfortable with hospital algorithms? Quietly judging our behavior to decide how loud our genetic voice gets to be. That is a fascinating question to leave on. This episode was based on an open access article under the CCBY 4.0 license.
23:37You can find a direct link to the paper and the license in our episode description. If you enjoyed this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description.
23:49Now stay with us for an original track created especially for this episode and inspired by the article you've just heard about. Thanks for listening, and join us next time as we explore more science base by base.