Protocol note 06
Sampling uncertainty is not representativeness
Separate random sampling error from coverage, nonresponse, measurement, and selection bias.
A narrow margin of error describes one modeled source of uncertainty under stated assumptions. It does not certify that the recruitment frame covers the target population, respondents resemble nonrespondents, questions measure the intended concept, or an opt-in sample is representative.
Distinguish four error families
Sampling error is the variation created by observing a sample instead of the full frame under a defined probability design. Coverage error arises when the frame omits, duplicates, or wrongly includes units. Nonresponse error appears when the people who do not participate differ on relevant outcomes from those who do. Measurement and processing errors include misunderstood questions, mode effects, coding mistakes, duplicate handling, and incorrect transformations.
A conventional margin of error addresses only sampling variation under its assumptions. It does not incorporate the other error sources, and it says nothing about whether a message study measured the intended interpretation. Publish the numeric interval and the non-sampling risks as different statements.
Probability and nonprobability samples support different claims
In a probability sample, each frame unit has a known nonzero selection probability and the design supports design-based uncertainty calculations. Even then, poor frame coverage, nonresponse, measurement, or weighting can weaken inference beyond the reported sampling interval.
An opt-in panel, convenience sample, river sample, or open website poll does not become probability-representative because quotas resemble census margins. Model-based intervals may be useful when their assumptions are explicit, but they should not be labeled as the probability-sample margin of error. Describe the recruitment mechanism and restrict conclusions to what the method can support.
Weighting can align observed variables to targets. It cannot guarantee alignment on unobserved variables related to both participation and the outcome. Publish the variables, source targets, trimming or calibration choices, effective sample size, and sensitivity checks rather than presenting weighting as a cure.
Worked example: large sample, unresolved bias
Study A draws 400 adults from a current address frame with documented selection probabilities but experiences substantial nonresponse. Study B recruits 20,000 volunteers through a product newsletter. Study B will usually have a much smaller simple-random-sample standard error if that formula is applied mechanically, yet its volunteer and customer-only selection may be strongly related to brand familiarity and message reaction.
The correct conclusion is not that Study A is automatically representative or that Study B is useless. Study A should report its frame, disposition, response, weighting, and nonresponse assessment. Study B can still compare randomized message variants within its recruited sample when implementation is sound, but population prevalence claims require stronger assumptions and cautious language.
A very large sample can make a biased estimate look stable. The discrepancy between a sample estimate and its target can grow with the interaction of population size, data quality, and quantity; more data do not erase systematic selection.
Choose a reporting lane before seeing results
Use a probability-population lane only when the frame, selection probabilities, and estimation method support it. Use a model-assisted lane when the model and calibration assumptions are specified, diagnosed, and made available. Use a sample-descriptive lane for opt-in or convenience recruitment when broader population inference cannot be justified.
A sample-descriptive statement can still be decision-useful: “Among the 612 eligible opt-in panel participants who completed this experiment, random assignment produced an 8-point difference between versions under the stated analysis.” It should not silently become “U.S. adults prefer version B.” Report uncertainty for the randomized contrast and separately limit the population scope.
Public-results checklist and limits
Every release should let a reader locate the boundary of the inference:
- Target population, frame population, sampling or recruitment source, mode, geography, language, and dates.
- Selection method, incentives, invitations when known, achieved completes, exclusions, and duplicate controls.
- Weighting variables and targets, design effect, effective base, missing-data handling, and uncertainty method.
- Exact question wording, stimuli, order, coding, outcome definitions, and all material deviations from plan.
- A plain statement of who and what the result does not represent.
Sources and scope
- Standards and Guidelines for Statistical SurveysU.S. Office of Management and Budget
Federal standards distinguishing sampling, coverage, nonresponse, measurement, processing, and model-related error sources.
- Sampling and non-sampling errorsStatistics Canada
Official educational reference separating sampling variability from coverage, nonresponse, response, and processing errors.
- Disclosure StandardsAmerican Association for Public Opinion Research
Professional disclosure expectations for sample source, recruitment, weighting, precision, question wording, and limitations.
- Statistical Paradises and Paradoxes in Big DataThe Annals of Applied Statistics
Peer-reviewed framework showing why data quantity cannot compensate automatically for selection-related data-quality problems.
Sources support the specific statements described above; they do not validate this publication’s rubric, guarantee a compliant execution, or replace context-specific professional advice.
What this page is: a research-methods guide, not a report of completed consumer fieldwork. Corrections and material revisions are recorded under the publication’s editorial standards.