A concise breakdown of a PNAS perspective that analyzes forensic proficiency testing practices using the 2023 CTS firearms test as a case study. The authors identify test design and administration flaws—easy items, consensus scoring, handling of inconclusives, nonblind verification, shot‑to‑shot variability, and contextual bias—that undermine claims about examiner accuracy and the utility of reported error rates for courts and laboratories.
0:00Welcome to Base by Base, the paper cast that brings genomics to you wherever you are. Thanks for listening, and don't forget to follow and rate us in your podcast app. So, uh, we usually focus strictly on genomics here, but today we are taking a deep dive into a slightly different kind of physical matching system, one that has equally massive consequences for human lives.
0:21Yeah, it's a really fascinating pivot I think. Right. So I want you to imagine for a 2nd that you are sitting on a jury. You know, you are in this quiet, really tense courtroom, and a highly credentialed forensic expert takes the stand.
0:35Getting a scene. Exactly. And they look right at the jury box and state with like 99% certainty that the bullet recovered from the crime scene was fired by the defendant's gun, not a gun that looks like it.
0:47That specific gun. And that is a powerful moment for a jury. Oh, absolutely. It's the kind of definitive statement that just locks up a conviction. And as a juror, you'd probably think, well, I mean, they're the expert that took a test to prove they could do this, right?
1:00You certainly hope so. But what if the test that certified that expert skill in the 1st place was, well, structurally flawed? What if the science wasn't actually as solid as the absolute confidence they're projecting on the stand?
1:13I mean, it is a genuinely terrifying thought to consider. Because on one hand, the human visual system is, well, it's genuinely amazing at classifying objects, right? Like we can instantly recognize a friend's face in a crowded room or, you know, distinguish a $5 bill from a $one bill at a glance.
1:33We do it so effortlessly that we just trust our eyes implicitly. Yeah, we don't even think about it. Exactly. But applying that everyday visual measurement system to high stakes forensic identity claims, especially without rigorous empirical validation, is incredibly risky.
1:48You are essentially asking a human being to perform as this like flawless microscopic measuring device. Okay, let's unpack this. Because today's deep dive is frankly a little unsettling. Today, we celebrate the work of researchers Nicholas Skirich and Thomas D.
2:03Albright, who have advanced our understanding of forensic identification evidence. Yeah, their recent paper is a real wake up call. It really is. It's a 2026 paper published in the proceedings of the National Academy of Sciences.
2:14And our mission today is to figure out why a recent forensic firearms proficiency test revealed this just staggering 20% failure rate among certified examiners. And uh, what that failure rate actually tells us about the hidden systemic weaknesses in forensic science?
2:31Right. And what actually happens when that flawed science walks into a real courtroom? So to really grasp what Scourage and Albright are arguing here, we have to look closely at this concrete stress test that recently shook the forensic community.
2:45Yeah, so this was a 2023 proficiency test. administered by an independent testing agency. And these agencies provide these exams to forensic labs so they can maintain their accreditation. Okay, so this isn't just a pop quiz for fun.
2:58No, at all. This is the actual exam that lets these labs keep their doors open, and, you know, lets these experts keep testifying in court. Exactly. It's high stakes for the lab. So before we get to that failure rate, how does a test like this actually work physically, are they just like holding 2 bullets up to the light?
3:13Uh not exactly. They use comparison microscope. So when a gun is manufactured. I mean, the machining process leaves these microscopic imperfections inside the steel barrels. Like wheel scratches. Right.
3:26And when you fire a gun, the explosive pressure forces a slightly softer metal bullet. So usually lead or copper straight down that steel barrel. And the intense friction scrapes the bullet, leaving microscopic scratches or strayations along its sides.
3:40Okay, making sense so far. So the working theory is that every single gun barrel has this entirely unique landscape of imperfection. So, it should leave a totally unique barcode of scratches on the bullet.
3:52So the examiner puts a crime scene bullet under one side of a, like a split screen microscope and a test fired bullet from the suspect's gun under the other side, and they just look to see if the barcodes line up.
4:02Precisely. core methodology. So in this specific 2023 proficiency test, the testing agency sent the examiners a physical kit. Like a box in the mail. Yeah, exactly. And inside, they were given what they call item one, which is a known bullet.
4:16The examiners know exactly which gun fired it. Then they were given 4 questioned bullets. Items two, three, four, and five. And the task was to examine those microscopic striations and determine if any of those 4 question bullets were fired by the exact same gun that fired item one.
4:35Right. So if I'm the examiner taking this test, I'm probably expecting at least one of these mystery bullets to match, right? I mean, that's how tests usually work. What was the actual ground truth of this specific test?
4:46The ground truth was that absolutely none of the question bullets were fired by the gun that fired item won. Wait, none of them. none of them. It was essentially a trap. Items two, 3, and 5 were fired by a 2nd gun, and item 4 was fired by a completely different type of gun altogether.
5:01Meaning the only factually correct definitive answer for every single one of those bullets is, no, these do not match. Exactly. So every single time an examiner looked at one of those bullets and confidently declared, yes, this is a match.
5:13They were committing a false positive error. They were definitively linking a bullet to a gun that never touched it. And the results were just staggering. If we exclude item 4 for a moment, and we'll get into the mechanics of why we're doing that in a bit, the false positive error rate for the remaining bullets was 19.2%.
5:31I honestly just need to pause and put that number into perspective for you, the listener, because 19.2% is astronomical in a high stakes field. It really is. The paper brings up this incredible comparison that really drove it home for me.
5:46Like, the FDA will halt the use of a life-saving medical device if there is a 0.02% chance of death. Right, a tiny fraction of a percent. That is almost a 1000 times less than this error rate. I mean, if you were running a cookie factory and you had a defective scanner on the assembly line, with a 20% false positive rate, you know, throwing out one out of every 5 perfectly good cookies.
6:10You would go out of business. You would shut the machine down. It would be a total financial disaster. But here we were talking about a tool used to put human beings in a prison cell for the rest of their lives.
6:19And you would think, right, that a nearly 20% failure rate would prompt this massive internal audit of the entire field. You'd hope so. Well, the leading professional organization for firearm examiners did launch a committee to investigate the fallout from this test.
6:34But their conclusion wasn't that the foundational science was flawed. Let me guess. They blamed the individual examiners. So they basically just went with the a few bad apples defense. They did. Yeah, they claimed that the examiners who failed were, you know, letting task irrelevant information influence them or that they maybe misunderstood their own laboratory's policies on how to report results, or they just flat out misjudged the visual agreement between the marks.
7:02I just I don't buy that for a second. I mean, if one person in a class fails a test. Sure, maybe they didn't study. But if 20% of your fully certified practicing experts fail it, you don't have a bad student problem.
7:14You have bad test problem. Or a bad curriculum. Which is exactly Scoorich and Albright's argument in the paper. By hyper focusing on the individual failures, the industry completely sidesteps the massive structural flaws in the test itself.
7:29And to understand how deep those flaws go, we need to talk about why we excluded item 4 from that initial error rate calculation. Okay, yeah, you mentioned item 4 was fired for a completely different gun.
7:41But aren't these proficiency tests supposed to simulate the hardest possible crime scene scenarios, you know, to really prove these experts know their stuff? You would hope so, but historically, the testing agencies design these proficiency tests to be, and this is by their own admission.
7:56They are designed to be relatively easy for a competent analyst pass. Oh, wow. Yeah. And item 4 is a perfect example of this. It was what the industry calls an out of class elimination. Walking through the physics of that.
8:07What does out of class actually look like under the microscope? So remember how we talked about the barrel leaving a microscopic barcode of scratches? Well, before you even get down to the microscopic level, you have the macro level design of the barrel itself.
8:20This is called the class characteristic. A conventional gun barrel is machined with sharp, distinct grooves cut into it. They kind of look like little staircases spinning down the inside of the barrel.
8:31Okay, got a visual of that. But a completely different class of gun might use polygonal rifling, which doesn't have sharp cuts at all. It's shaped much more like a smooth, twisting octagon. So if the known bullet was fired from a gun with sharp grooves, and this item 4 bullet was fired from a gun with a smooth octagon barrel.
8:53I mean you don't even need a microscope for that, do you? Not really, no. You could probably just look at it with a magnifying glass and say, these are entirely different shapes. Exactly. Comparing them requires almost no subjective interpretation.
9:04And the paper uses this really fantastic analogy here. I love a good analogy. Right. So including an out of class elimination on a high-level proficiency exam for forensic scientist is basically like asking a professional race car driver on their licensing exam what a red stoplight signifies.
9:21Right. It doesn't tell you who the best drivers are at all, it just proves they aren't completely oblivious to the most basic rules of the road. Precisely. But if it's that obvious, surely nobody failed item four, right?
9:33Well, that is actually the most alarming part of the data, even on this incredibly basic question. one.5% of examiners failed. Wait, really? Yeah, and one examiner actually called it a positive identification.
9:47Hold on. They said, a bullet with smooth octagonal marks perfectly matched a gun with sharp staircase crews. Yeah. But that's physically impossible. It is completely physically impossible. And even worse, 4 other examiners looked at that same glaringly obvious difference and marked the result as inconclusive.
10:06How do you call a red stoplight inconclusive? Do they just say, well, um, it's a shade of light. I can't really be sure. The wild part is that the industry actually defended those inconclusive responses.
10:17No, their argument was that if the difference between the conventional and polygonal rifling wasn't discernible to the examiner, then, well, calling it inconclusive is technically the appropriate, cautious response for that specific examiner.
10:33Wow. So therefore, it shouldn't be counted as an error. That is just maddening. I mean, if I take a vision test at the DMV, and I tell them, you know, the giant E on the chart is inconclusive to me, they don't say, oh, good job being cautious.
10:47Here's your license. Right. They tell you can't drive. Exactly. They tell me my vision is terrible. And that is exactly Skeerage and Albright's point here. This points to a fundamental failure of examiner training in basic class characteristics or, you know, a failure in their actual visual sensitivity.
11:03It is not some magical, ambiguous property of the bullets themselves. If 100s of other examiners could easily see the physical difference, excusing the few who couldn't, basically allows the field to avoid asking a very uncomfortable question.
11:18Which is why can't some of our certified experts pass a basic baseline check? And if examiners are hiding behind inconclusive on a question as obvious as a red light. It really makes me wonder what they do when the evidence is actually difficult to read.
11:34Like, do they just guess or do they use inconclusive as this giant safety net? It is absolutely used as a safety net. And this brings us to the most controversial part of how these proficiency tests are actually graded, because across this entire 2023 test, nearly half 48% of the total responses were marked inconclusive.
11:54Half the test. half the test. So how do you even grade an exam where half the answers are essentially just, I don't know, because if the ground truth is that none of the bullets matched, then saying, I don't know is wrong, the definitive answer is no match.
12:07Logically, yes, you're totally right. If you use ground truth as your standard, like the actual physical reality of where those bullets came from, every single inconclusive response on this test is an error.
12:18But the testing agency doesn't grade based strictly on ground truth. They use a system called consensus scoring. Wait, consensus? You mean they just look at what the majority of people answer? Yes, exactly.
12:29Whatever the majority of examiners guess. basically becomes the correct answer for the grading key. Oh my god. And since the vast majority of examiners on this test guest inconclusive, the testing agency actually scored inconclusive as a correct response.
12:45Hold on, I need to make sure I'm fully understanding this. You're saying that in what claims to be a hard science, if 51% of the experts get the answer mathematically, physically wrong, the testing agency just updates the answer key and marks the wrong answer as correct.
13:01That's how it works. How can they legally justify that in a courtroom? I mean, and I take a calculus exam and the whole class gets the equation wrong. We don't all get an A just because we agreed. You hit on the core absurdity of it, really.
13:13Consensus scoring is a tool that's usually reserved for fields without objective ground truth. Like what? Think about, like, an art appraisal or judging a gymnastics routine. You need a panel of experts to reach a consensus because there is no strict mathematical answer.
13:28Okay, that makes sense for art. Right. But in a field claiming to be a hard science based on physical mechanics, using consensus scoring creates this totally dangerous feedback loop. If the majority of the field is poorly trained or using a flawed methodology, the consensus will inherently be flawed, but the grading system will validate them anyway.
13:48It's just an echo chamber of error. So how do we fix it? Like does the paper offer a way out of this grading nightmare? They do, yeah. And it requires throwing out that vague, inconclusive bucket entirely.
14:01Skirish and Albright suggests implementing something called signal detection theory. Okay, what does signal detection theory actually look like in practice? Let's say I'm a forensic examiner sitting at my desk on a Tuesday morning.
14:12How does my job change? So right now, you essentially have 3 bins to drop your answer into. Yes, no, or inconclusive. It's a really blunt instrument. Under signal detection theory, instead of those rigid bins, you would have to provide a continuous confidence rating.
14:27Say, a scale from one to 10. One means you are absolutely certain it's not a match. 10 means you're absolutely certain it is, and 5 means you genuinely can't lean either way. Ah, so it's like a slider rather than just on off switch.
14:40Exactly. And over 100s of tests, a lab can map those confidence ratings on a curve to see an examiner's true accuracy. That smart. Yeah, you plot their sensitivity to true signals, meaning how often they confidently catch an actual match against their susceptibility to noise, which is how often they confidently call a false match.
14:59Okay, I fell. It allows labs to mathematically measure discriminability, which is their actual ability to separate reality from illusion rather than just tallying up majority votes to make everyone. look good.
15:10That sounds like actual rigorous science. But, um, Even if we perfectly fix the grading scale, it assumes the test itself is physical sound, right? Like if I'm grading on a perfect curve, but my exam paper is literally printed with invisible ink.
15:26The grade still doesn't mean anything. That's a great point. Which makes me wonder, are the actual physical bullets they send out in these kits consistent? That is the multimillion dollar question, and it leads us to this massive bombshell, hidden deep in the industry's own investigative report.
15:42get laid on me. So one laboratory decided to buy 2 separate test sets from the testing agency for their examiners to use. Now, keep in mind these sets are manufactured to be completely identical. They're supposed to be standardized tests.
15:57Let me guess, they weren't identical. They were drastically different. The bullets in the 1st set were clearly marked with deep readable striations that allowed the examiner to actually reach a definitive conclusion.
16:11But the bullets in the 2nd set were a blurry, poorly marked mess. Yikes The report noted that for the 2nd set, inconclusive was the only scientifically appropriate response, because the physical data just wasn't even there.
16:24But wait, I thought a gun barrel is made of solid steel. It doesn't just change its shape from shot to shot, does it? If you fire 10 bullets from the same gun, shouldn't they all look exactly the same?
16:35Well, that is the foundational myth of firearm identification. In reality, a gun barrel is an incredibly hostile dynamic environment. Right. You have extreme heat, explosive pressure, friction, metal expanding and contracting, lead residue building up, unburned powder acting like sandpaper.
16:52Sounds chaotic. It is, the physical conditions inside the barrel change milliseconds after every single shot. So the 100th bullet, fired from a gun, might have a significantly different microscopic landscape than the 1st bullet fired.
17:06Which means we are dealing with a massive reproducibility crisis. If the testing agency who is, you know, deliberately trying to manufacture perfect identical test bullets in a controlled laboratory. If they can't even get 2 bullets from the same gun to look the same.
17:21What does that mean for a random gun fired in an alleyway? It means we really have to question the core premise of the entire field. Because if you were an examiner and the police bring you a single battered crime scene bullet, but they haven't found the suspect's gun yet.
17:34How do you evaluate those scratches? How do you know if a specific mark on that bullet is a permanent, unique characteristic of that specific barrel? You don't. I mean, it could just be a random transient imperfection.
17:47Like a grain of sand could have been in the barrel for exactly one shot, scratch the bullet, and then blown out. You could just be matching random noise. Precisely. Without the actual firearm in hand, to extensively test fire and confirm that a specific pattern of marks is reliably reproducible shot after shot, linking a single piece of copper to a unique source in the world becomes incredibly tenuous.
18:09And this brings us to a really honestly dark realization about the human mind. Because if the physical evidence is this variable, and often just looks like a blurry mess. What are the examiners actually looking at?
18:24Well, that's where psychology comes in. Right. If the physical data isn't there, their brains are going to try to fill in the blanks. And where there is subjective visual judgment, cognitive bias is going to creep in.
18:33Absolutely. Our visual system doesn't operate in a vacuum, expectations heavily, heavily shape our reality. think about the proficiency tests themselves. When an examiner opens that kit, they naturally fall into the trap of assuming it's a closed set.
18:48They subconsciously expect that the testing agency wouldn't just send them a puzzle without an answer, so they assume at least one of the question bullets must be a match to the sample. Right. It's like a test-taking bias.
19:00If you assume there's a pattern to find, your brain will literally invent a pattern in the static just to solve the puzzle. Yes, and it is so much worse in actual real world case work. How so? In a real investigation, examiners aren't just given bullets in a blank envelope.
19:16They are often given highly suggestive contextual information. The paper highlights how investigative leads are generated by the national ballistic database, right? Nibin, I think is called. Yes, exactly.
19:28And when the computer algorithm flags a potential connection between 2 crimes, The printout handed to the human examiner explicitly instructs them to confirm the match. It literally says the word confirmed.
19:39It literally says confirm. That is staggering. It completely anchors their expectation before they even look in the microscope. You aren't being asked to analyze the evidence objectively to see if it matches.
19:50You're being asked by law enforcement to rubber stamp a computer's hypothesis. The psychological pressure is immense. And it completely compromises the way labs try to catch errors, specifically their verification process.
20:04How does that work? Well, when an examiner makes a positive identification. Standard lab procedure requires a 2nd examiner to verify that conclusion. Okay, that sounds good in theory. Like peer review.
20:17In theory, yes. But the paper points out that this verification is almost always non-blind, meaning the 2nd examiner knows exactly what the 1st examiner concluded before they ever look into the microscope.
20:28Oh, come on. Having a 2nd examiner verify your work while staring at your written conclusion is like, it's like proofreading an essay where all the autocorrect changes have already been accepted. You aren't actually reading the words anymore.
20:40You just see that the green check mark in your brain goes, looks good to me. It's not science. It's just a social compliance test. That is a perfect analogy. You are verifying the social expectation, not the physical evidence.
20:52And the empirical data backs this up entirely. Really? Yeah, the paper cites a European study that looked at this exact phenomenon. They found that examiners were 5 times more likely to disagree with the initial conclusion if they were truly blind to it, meaning they had no idea what the 1st examiner decided.
21:08Five times. And the bias was even stronger if the 1st examiner had a higher professional status or seniority in the lab. So non-blind verification doesn't catch errors. It mathematically reinforces them.
21:21It creates this closed loop that gives the false impression of immense scientific rigor. Wow. This has been an incredibly heavy deep dive and honestly a bit infuriating. Let's pull all these threads together for you listening right now.
21:33What we're seeing is that a 20% false positive rate on a proficiency test isn't just about a few examiners having a bad day or, you know, needing new glasses. at all. It is a blazing neon warning sign about deep structural flaws.
21:47The tests are designed to be artificially easy. They use consensus scoring to hide how often experts just guess inconclusive, the physical bullets themselves suffer from massive shot to shot variability, and the examiners are just swimming in a soup of cognitive bias and peer pressure.
22:04And if you're wondering why this matters to you outside of an academic debate, it's because these exact flawed proficiency tests are frequently the bedrock of courtroom testimony. Right. When judges are trying to decide what science is reliable enough to show to a jury, they apply legal frameworks known as the daubert or fry admissibility standards.
22:24Think of these standards like, um, like a bouncer at a nightclub designed to keep junk science out of the courtroom. But the bouncers being handed a fake ID. I mean, prosecutors point to these proficiency test, with their consensus scored artificially inflated pass rates, to claim that forensic firearm identification is nearly infallible.
22:41They tell juries it has a 0 or one% error rate. Exactly. They use these tests to completely bypass the legal gatekeepers, and presenting this overstated mathematically false certainty to a jury poses a massive risk of wrongful convictions. Juries believe it because it's presented as hard physical science.
23:01and Scarage and Albright make it abundantly clear in the paper. The field desperately needs a cultural and structural shift. We need radical transparency in how these tests are graded. We need to move away from subjective human eyeballs, and toward objective 3D surface scanning algorithms, and labs absolutely must mandate truly blind verification protocols.
23:23Absolutely. I want to leave you with a final thought to mull over today. It builds on something the paper touched on. Viewing these proficiency tests as a calibration check. I like that, framing. Think about a police breathalyzer, right?
23:35Or a radar gun used by a highway patrol officer. If the police department took that radar gun into the shop, ran a calibration stress test, and found that 20% of the time the machine was completely wrong.
23:45They would throw the machine in the garbage. They wouldn't trust a single ticket it produced. Exactly. But in forensic firearms analysis, the machine taking the measurement is a human being's visual judgment.
23:57If this 20% failure rate is the true unvarnished state of our current calibration. What does that mean for the 1000s of closed cases, and the real people currently sitting in our prison system, whose fates were decided before anyone even realized the scale was broken?
24:12It is a profound and, frankly, unsettling question that the justice system can no longer afford to ignore. This episode was based on an open access article under the CCBY 4.0 license. You can find a direct link to the paper and the license in our episode description.
24:28If you enjoy this, follow or subscribe in your podcast app and leave a 5 star rating. If you'd like to support our work, use the donation link in the description. Now, stay with us for an original track created especially for this episode and inspired by the article you've just heard about.
24:41Thanks for listening, and join us next time as we explore more science, base by base.