At some point in your PCI DSS assessment, your QSA is going to tell you they didn’t test everything, and that’s normal. No assessment tests every system, document, or physical location in your environment. What surprises a lot of companies is how much method sits between “we didn’t test all of it” and “we tested enough of it to mean something.”
What Does “Sampling” Mean in a PCI DSS Assessment?
Sampling is the practice of testing a representative subset of an environment, rather than every single system, document, or location, to reach a defensible conclusion about the whole. PCI DSS assessments rely on sampling because most environments are too large to test in full within a reasonable engagement timeline, and PCI SSC provides assessors with a structured methodology for how large a sample needs to be relative to the total population it’s drawn from.
Sampling, briefly
PCI DSS assessments run on sampling. The process starts by categorizing your environment, system components, individuals, documentation, by type and function, then sizing a representative sample from each category. A population of four to thirty components might only require a sample of three. A population over a thousand still caps at a sample of thirty. Physical locations follow a related but separate logic, tied to whether security policy is set centrally or per location.
What that looks like
In practice, a sample gets built and sized against the population it’s drawn from. A hundred web servers running the same configuration might only need ten tested. Three domain controllers, all functionally identical, might mean testing all three, since the population is already too small to sample further. A single bastion host gets tested on its own, there’s nothing to sample it against. None of this is arbitrary: it’s a structured method for turning a population too large to fully test into a defensible, representative slice of it.
Sampling for Individuals and Documentation, Not Just Systems
System components aren’t the only thing that gets sampled. PCI DSS assessments also sample individuals, employees interviewed about security awareness training, for instance, and documentation, like firewall change records, using the same underlying logic: categorize first, then size a sample against the population. A population of security awareness training records numbering in the thousands might still cap at a sample of thirty, the same ceiling that applies to system components at that scale. The method stays consistent even though what’s being sampled changes.
Physical Locations Follow a Different Rule
Physical security controls get sampled differently than systems or documents. Whether a location needs to be sampled or visited individually depends on how its security policy is governed. A single corporate headquarters with its own dedicated policy typically gets tested on its own. Retail locations or call centers governed by a shared regional or central policy can be sampled at a much smaller ratio, sometimes as few as three locations representing hundreds, because the assumption is that a centrally governed policy doesn’t vary meaningfully from site to site. That assumption is only safe if it’s actually true, which is itself worth confirming rather than assuming.
Why Categorization Happens Before Sizing, Not After
The order of operations matters more than it sounds like it should. A population has to be categorized by type and function before a sample size gets calculated, not the other way around. Categorize too broadly, lumping a web server farm and a set of administrative workstations into one bucket because they’re both “servers,” and the resulting sample undercounts a population that was actually more diverse than the category suggested. Categorize accurately, and the sample size formula does what it’s supposed to: produce a slice of the environment that’s actually representative of what’s really there.
Why the assessor matters here, specifically
The math behind sample sizing isn’t where a rushed assessment and a thorough one diverge, both can follow the same formula and land on the same number. The divergence shows up earlier, in how carefully your environment gets categorized in the first place. An assessor who lumps dissimilar systems into one category to shrink the required sample is technically following the methodology and still missing coverage that actually matters.
Where quality control comes in
This is also where PCI SSC’s own rules add a layer most companies never hear about: quality assurance reviews of an assessment have to be completed by separately qualified personnel, not just the assessor who did the fieldwork. A QSA firm that takes this seriously has someone checking sampling and workpapers before a report goes out, not after a customer or acquirer questions it. Ask a vendor how their internal review actually works and you’ll usually learn more about their standards than any credential alone would tell you.
What Happens When a Sampled Item Fails Testing?
If a control fails testing on even one item in a sample, the finding isn’t treated as isolated to that single system or record. The assessor generally has to consider whether the failure indicates a broader gap across the rest of the population it was drawn from, which can mean expanding the sample, retesting a wider set once the underlying issue is remediated, or, in some cases, treating the whole category as non-compliant until the fix is verified more broadly. A single failed item in a sample of thirty can end up requiring far more remediation work than the ratio suggests, precisely because the sample is meant to represent everything it was drawn from, not just the piece that happened to get tested. That’s also why a rushed retest, one that only re-checks the single failed item instead of the broader population it represents, can leave a report technically passing while the underlying gap is still there.
What to ask a vendor
- How do you categorize system components before sizing a sample, and who reviews that categorization?
- Does someone other than the assessor who did the fieldwork review the sampling and workpapers before the report is issued?
- If our environment includes multiple locations, how does sampling work across them?
PCI DSS gives assessors real judgment in how sampling gets applied, so there’s no single right answer to grade these against. What they get at is whether that judgment is being exercised carefully, or just quickly.
Common Questions About PCI DSS Sampling
Does a PCI DSS assessor test every system in scope? No. PCI DSS assessments use sampling, testing a representative subset of in-scope systems, documents, and locations rather than the entire population, following a structured methodology tied to population size.
Who reviews an assessor’s sampling decisions? PCI SSC requires QSA Companies to have an internal quality assurance process, meaning a separately qualified reviewer, not just the assessor who did the fieldwork, checks sampling and workpapers before a report is issued.
Can a smaller sample mean a less thorough assessment? Not by itself. Sample size follows a defined methodology based on population size. What varies between assessors is how carefully the environment gets categorized before that sample is drawn, which affects how representative the sample actually is.
Contact us to talk through how we’d approach sampling for your environment, or learn more about how Insight Assurance runs PCI DSS assessments.

