Protocol note 01
A practical protocol for testing sponsored-content disclosures
Compare disclosure labels without confusing recognition, comprehension, and trust.
A disclosure can be visible without being understood. This protocol separates three questions that are often collapsed: did people notice the label, did they understand the commercial relationship, and did that understanding change their interpretation?
Define the decision before the questionnaire
Start narrowly: which eligible label most clearly communicates that the named brand paid for the article? That is testable. “Which disclosure is best?” is not, because best could mean noticeable, accurately understood, trusted, familiar, or unobtrusive.
Pre-specify one primary outcome before recruitment. Accurate unaided description of the commercial relationship is a defensible choice. Unaided recall, aided recognition, confidence, perceived trustworthiness, reading intent, and incorrect beliefs about sponsor control can remain separate secondary outcomes.
A decline in trust is not automatically a disclosure failure. Correct recognition that material is advertising can change how readers judge it. Choosing the label with the smallest trust or click penalty could reward ambiguity rather than understanding.
Build comparable conditions
For a language test, keep placement, contrast, typography, headline, image, article copy, sponsor, page width, scroll position, and exposure rules identical. Change only the disclosure text. Test placement in a separate experiment so any difference has a clearer interpretation.
Randomly assign participants before exposure and retain the assigned condition in the study record. Define eligibility, quotas, completion rules, duplicate protections, quality screens, exclusions, and the minimum difference worth acting on before outcomes are visible.
Every candidate should first pass a policy and factual review. Do not deploy a label the publisher already believes is misleading merely to create a convenient control. A diagnostic study can reveal misunderstanding; it does not transform an inadequate disclosure into an acceptable one.
Use a question order that does not teach the answer
Ask broad questions before showing language that tells participants what they were supposed to notice. Put demographics, category-use questions, and hypothesis-revealing prompts later unless they are genuinely required for eligibility.
Use the same fictional or rights-cleared stimulus in every condition. The example below uses a fictional company, Northstar Forms, so the protocol itself does not imply a real commercial relationship.
- Unaided meaning: “What, if anything, did the label at the top of the page tell you?”
- Unaided relationship: “In your own words, what relationship, if any, exists between the publisher and Northstar Forms?”
- Aided interpretation: choose among paid placement, information supplied without payment, no brand involvement, or cannot tell.
- Notice and recall: ask whether a commercial label was noticed, then request the exact words remembered.
- Confidence: use a defined scale for confidence in the relationship description.
- Trust: ask about trust in the presentation independently of agreement with the article.
- Reasoning probe: “What on the page most influenced your answer?”
Predefine coding and analysis
Code the two open answers with a written scheme: accurate, partial, incorrect, or uncertain/blank. “Accurate” should identify the relevant payment or commissioned-content relationship; “partial” may recognize advertising without explaining the relationship. Coders should not see the assigned condition.
Use at least two coders for a documented agreement check when resources permit, and define how disagreements will be resolved. The primary comparison may reduce the outcome to accurate versus all other responses, but the report should still show partial, incorrect, and uncertain responses separately.
Report estimates with uncertainty rather than declaring a winner from raw percentages alone. Any weighting, exclusions, subgroup analysis, or multiple-comparison adjustment should be specified before the results are examined.
Worked example: three labels, one relationship
Assume the fictional Northstar Forms paid for an article while the publisher retained control of its analysis and conclusions. Three otherwise identical mobile pages place one label above the headline: “Advertisement”; “Paid advertisement from Northstar Forms”; or “Northstar Forms paid for this article. The publisher controlled the analysis and conclusions.”
The decision is not which label generates the most clicks. It is whether one produces meaningfully more accurate unaided descriptions without creating a false belief about editorial control. If the longer version improves only aided recognition, that is weaker evidence than improvement in the pre-specified unaided measure.
If the planned comparison does not distinguish the conditions at the stated precision, report that uncertainty. Do not quietly select the numerically highest label or redefine success after seeing the answers.
Failure modes to catch before launch
A clean-looking questionnaire can still bias the result. Review the full path against these common failures:
- Naming “the sponsorship disclosure” before measuring unaided understanding.
- Changing wording, color, placement, and page layout in the same comparison.
- Using label familiarity or visual attention as a substitute for relationship comprehension.
- Selecting the version with the smallest trust penalty instead of the clearest meaning.
- Reporting correct answers while hiding uncertainty, partial understanding, or specific wrong beliefs.
- Pooling countries, languages, devices, or recruitment sources that produced materially different exposure.
- Allowing a sponsor to suppress valid unfavorable findings.
- Treating a mock-page study as proof of legal compliance or live-market behavior.
Report the misses and the study’s limits
Publish the exact screenshots and label text, target population, recruitment source, mode, field dates, geography, language, achieved sample, assignment method, question order, exclusions, coding rules, weighting, sponsor role, and estimates with uncertainty. Show the share who misunderstood, gave a partial answer, said they did not know, or never recalled a label.
A simulated exposure cannot reproduce every distraction, scroll pattern, device setting, or prior brand belief. Self-reported notice is imperfect, and an opt-in sample should not be described as probability-representative without an appropriate design.
Consumer-interpretation evidence can identify risk; it does not itself establish compliance. This is a design proposal. Message Test Bench has not fielded this protocol or published empirical results from it.
Sources and scope
- Native Advertising: A Guide for BusinessesU.S. Federal Trade Commission
Guidance on clear disclosure language, prominence, proximity, and disclosure before engagement.
- Blurred Lines: Consumers’ Advertising RecognitionU.S. Federal Trade Commission staff
Exploratory usability research showing why visual attention and accurate recognition should not be treated as the same outcome.
- Disclosure StandardsAmerican Association for Public Opinion Research
Professional reporting standard for publishing enough study detail to permit independent evaluation.
Sources support the specific statements described above; they do not validate this publication’s rubric, guarantee a compliant execution, or replace context-specific professional advice.
Editorial status: This is an original research-methods note, not a report of completed consumer research. Corrections and material revisions are recorded under the publication’s editorial standards.