Skip to content

Why labs disagree: lab variance and 'lab shopping', with the data

Washington test data show cannabis labs systematically reporting higher THC than others for identical cultivars. The real numbers, and what they can't prove.

This depends on where you live. Plant limits, licensing, permitted products and testing rules differ by country and change often. Check the law section before acting on it.

On this page

Send the same Washington State flower category to six different accredited labs and, according to the only large peer-reviewed study of this question, the median THC result you get back can differ by up to 5.5 percentage points depending on which lab you pick [1]. That is not a rounding difference. On a 20% THC flower, it is the gap between a middling result and a headline number, and it survives every innocent explanation researchers could throw at it.

THC spread, chemotype I flower
17.6–23.1%
median across six WA labs, pooled 2014–2017, corrected 2020 figures [1]
Effect size, lowest vs highest lab
1.27Cohen's d
an 82% chance a random high-lab reading beats a random low-lab one [1]
99th-percentile THC, low- vs high-reporting lab
26.9 vs 31.7%
same flower category, different lab [1]
CBD, chemotype II/III flower
~13% (highest lab)
significantly above every other lab, p under 0.01 [1]

What the Washington data actually is, and what it isn't

The number above comes from one specific, well-documented source: Jikomes and Zoorob's 2018 analysis of Washington State's seed-to-sale traceability system, obtained by public records request and covering 215,285 cannabinoid test results from June 2014 to May 2017, across every state-licensed testing laboratory [1]. The same dataset shows reported THC rising through 2014-15 and then largely flattening; see how lab-reported THC has drifted over time for that trend on its own. It is worth being precise about what kind of evidence this is, because the word "lab variance" gets used loosely.

This is a market-level statistical study, not a blind proficiency test. The authors did not send one physical sample to six labs and compare the results; they compared what six labs reported, in aggregate, across the different flower and concentrate samples that growers and processors actually submitted to each of them over three years. The six largest labs by data volume, referred to as Labs A through F, are named in the paper: Confidence Analytics, Analytical 360, Green Grower Labs, Integrity Labs, Testing Technologies and Peak Analytics [1]. For chemotype I flower (high-THC, low-CBD, the dominant commercial type), median reported total THC ranged from 17.6% at the lowest-reporting lab to 23.1% at the highest [1].

The differences survive every obvious explanation

The first, generous explanation for a gap like this is that different labs simply see different products: one lab gets sent a lot of Chemdawg-type cultivars, another mostly imports from a different set of growers, and the "difference between labs" is really a difference between what each lab was handed. Jikomes and Zoorob tested this directly. They ran fixed-effects regressions that strip out the grower, the cultivar name and the test date, leaving only the lab-specific deviation, across four separate models: THC in chemotype I flower (n=115,626), THC in chemotype I concentrates (n=24,065), CBD in chemotype II and III flower (n=2,955), and CBD in the same chemotypes' concentrates (n=1,125) [1].

The gaps did not close. After adjustment, Peak Analytics' average chemotype I flower THC (about 23%) was still significantly higher than every other lab (p under 0.001), and the pattern repeated for concentrates (about 75% versus a next-highest 70%) and for CBD in chemotype II and III flower, where Peak Analytics' mean of about 13% stayed significantly higher than every other lab (p under 0.01) [1]. Because the same producers, the same cultivar names and the same weeks of testing were accounted for, the authors conclude the remaining gap is "likely caused by systematic differences in [labs'] testing methodologies," not by which growers happened to use which lab [1].

Effect sizes make the practical stakes clearer than the p-values do. Comparing the two closest labs (A and B) gives Cohen's d=0.11, a small effect: a random sample from the higher lab beats a random sample from the lower one only 53% of the time. Comparing the two extremes (A and F) gives d=1.27: an 82% chance the higher-reporting lab's random reading beats the lower-reporting lab's [1]. Labs also differ sharply in how often they report near-zero CBD in high-THC flower, plausibly a limit-of-quantification difference between methods, with effect sizes (Cohen's h) frequently exceeding 0.80 between lab pairs [1].

A useful sanity check sits in the tails. Restricted to the lowest-reporting lab, the 99th-percentile chemotype I flower result is 26.9% THC; restricted to the highest-reporting lab, it's 31.7% [1]. The authors' own reading of that gap: flower genuinely testing above 30% total THC is rare, and much of what gets labelled that way is probably not accurate [1]. Independent measurements from outside Washington back the low end, not the high one: Vergara et al. (2017), as reported within Jikomes and Zoorob's discussion, found commercial flower averaging around 19% THC in Seattle and closer to 15% in Denver, Sacramento and Oakland; a separate sample of Californian medical flower, cited the same way, averaged around 17% [1].

Scatter chart: Washington labs, chemotype I flowerWashington labs, chemotype I flower: 6 points, peak 60,924 n at 18.020,00040,00060,00080,000Chemotype I flower samples tested (n)1618202224Mean reported THC (%)Washington labs, chemotype I flower: 18 %, 60,924 nWashington labs, chemotype I flower: 18 %, 17,347 nWashington labs, chemotype I flower: 20 %, 31,160 nWashington labs, chemotype I flower: 20 %, 21,885 nWashington labs, chemotype I flower: 22 %, 13,372 nWashington labs, chemotype I flower: 23 %, 25,452 nLab ALab F
Fig. 1No clean line between a lab's reported THC and how much business it did in this dataset. The biggest lab by volume reported the lowest THC of the six. Values as originally published in 2018; a 2020 correction revised the underlying dataset (see Sources).Horus
LabMedian total THC, chemotype I flower (%)Chemotype I flower samples tested (n)
Confidence Analytics (Lab A)17.760,924
Analytical 360 (Lab B)18.417,347
Green Grower Labs (Lab C)19.731,160
Integrity Labs (Lab D)20.421,885
Testing Technologies (Lab E)22.113,372
Peak Analytics (Lab F)23.225,452

The lab-by-lab figures above are as first published in 2018. A 2020 author correction to the same paper removed duplicate database entries, which pulled the pooled dataset down to 215,285 test results and nudged the chemotype I flower THC range to 17.6–23.1%; it changed no lab's rank and did not alter the paper's conclusions [1].

The economics behind "lab shopping", and what this dataset does not show

"Lab shopping" is the name industry commentary gives to a simple incentive: if potency drives wholesale value and a producer can choose which accredited lab tests a batch, some producers will gravitate toward whichever lab reports the higher number. Jikomes and Zoorob's own discussion leans on exactly this logic. They cite a separate statistical analysis of the same Washington sales data that found a significant relationship between commercial flower value and its reported THC and CBD content, and argue this "may provide an economic incentive for cannabis producers to seek test results with higher total THC or CBD levels" [1]. They also note their findings are "consistent with reports of 'cannabinoid inflation' by certain laboratories" already circulating in industry press before their paper was published [1].

That is the mechanism, and it is a reasonable one. But it is worth testing against the one piece of volume data the paper actually publishes, rather than assuming it: each lab's total sample count over the study period, reported alongside its mean THC in the same figure [1]. Plotted against each other, they do not show a clean "report higher, get more business" line. Confidence Analytics (Lab A), the lowest-reporting lab of the six, tested by far the largest number of chemotype I flower samples, 60,924, nearly double the next-largest lab (Green Grower Labs, 31,160). Peak Analytics (Lab F), the highest-reporting lab, tested a middling 25,452, less than half of Lab A's volume [1].

That does not disprove a lab-shopping dynamic; the paper's own dataset is a three-year pool, not a lab-by-lab time series of submissions, so it cannot show whether any lab's market share moved up or down as its reported numbers moved. What it does show is that a large, established lab can hold the biggest share of the market while reporting the lowest numbers, which is a fact worth knowing before repeating the tidier version of this story. A correlation between a lab's average reported potency and how much of the market it holds, if one exists, is a signal about incentives across the whole system. It is not, on its own, evidence that any specific lab manipulated any specific sample; the authors are explicit that their data point to "systematic differences in testing methodologies" as the likely driver, a milder and more defensible claim than fraud [1].

Sarma et al. is a different paper, doing different work

It's common to see Sarma et al.'s 2020 paper cited alongside Jikomes and Zoorob as if it were a second variance study. It isn't. Sarma and co-authors, working with the US Pharmacopeia, published a monograph-style quality-attributes paper on cannabis inflorescence: hierarchical botanical nomenclature, macroscopic and microscopic identification, chromatographic methods (including HPTLC and HPLC versus GC) to identify and quantify individual cannabinoids rather than just total THC and CBD, limits for contaminants such as heavy metals, pesticides and microbial counts, and practical validation problems such as the instability of cannabinoid-acid reference standards used to calibrate an assay in the first place [2].

None of that is a measurement of how much labs disagree in practice. What it establishes instead is the methodological floor a "true potency" figure depends on: a validated reference standard, a defined and reproducible assay, and stated acceptance criteria. Cite it for that groundwork, not for a variance number it was never designed to produce.

Proficiency testing is the design that can actually isolate lab error

Jikomes and Zoorob's regression controls for grower, cultivar and date, which rules out the most obvious confound. What it cannot do is separate "Lab F's chromatography genuinely reads higher" from "growers happened to send Lab F their best-looking batches in ways the regression didn't capture." Only one study design closes that gap: a blind split sample, one physical batch of homogenised material, divided and sent simultaneously to multiple labs under a shared accreditation scheme, with none of the labs told the others' results or, ideally, that the sample is a test at all.

That design has a name: proficiency testing (PT), evaluated under ISO/IEC 17025, the international standard for testing and calibration laboratories. Clause 7.7 of the 2017 revision requires an accredited lab to monitor its own performance by comparison with other labs, specifically through participation in proficiency testing or other interlaboratory comparisons, and to take corrective action when its results fall outside pre-defined criteria [3]. Jikomes and Zoorob reach the same conclusion from the opposite direction: having found market-level inflation they cannot fully explain, they recommend states require third-party ISO/IEC 17025 accreditation for cannabis labs and run regular "round-robin" audits, sending a blinded, common sample across the labs operating in that market [1].

Reading a COA, and choosing a lab, with this in mind

A single high number is not proof of quality, and a single low number is not proof of honesty; both need context the certificate of analysis alone doesn't give you. Two things are worth checking before you act on a figure.

First, does the number sit inside that lab's own plausible range, or right at its outer edge? Using Jikomes and Zoorob's own tail data as a worked example: a chemotype I flower result of 28% THC from a lab whose 99th percentile normally sits at 26.9% is an outlier for that lab and worth querying. The same 28% from a lab whose 99th percentile normally runs to 31.7% is unusual but well inside what that lab's own method regularly produces [1]. The number means different things depending on which lab produced it, which is exactly why a single figure, without knowing the lab's own distribution, tells you less than it looks like it does.

Second, is the gap between two labs testing the same material a one-off or a pattern? A single disagreement can be batch variation, sampling error or genuine biological difference between two sub-samples of the same plant. A persistent gap between the same two labs across multiple batches is a testing-system signal, not proof of malpractice on either side, but a reason to ask both labs for their accreditation status and, where the volume justifies it, to enrol in a proficiency-testing scheme rather than simply picking whichever lab reports higher going forward.

What an inflated number costs once someone checks

The commercial exposure runs in one direction. A batch priced, or a compliance record filed, against a reported THC or CBD figure that a proficiency-testing round or a regulator's resample later contradicts does not just cost an argument with a lab, it can cost a contract price already paid on the higher number, or a compliance finding against a facility that relied on it in good faith, the same kind of paper trail that GMP inspections check for licensed operations. That exposure scales with volume in exactly the way a home grower's single-sample risk does not: one inflated certificate on a hundred kilograms of licensed flower is a real financial and compliance liability the moment an auditor or a buyer's own testing disagrees with it.

The fix is not to find the lab that reports the highest number and stay there. It's to treat any single lab's figure as one measurement from one method, ask what that lab's own track record against blind proficiency samples looks like, and build a resampling habit, at home a second opinion when a number looks unusually high, at scale a standing PT enrolment, before a pricing or compliance decision leans on a figure nobody outside that one lab has ever checked.