Copy-ready questionnaire
Message testing survey: 10 questions and a usable template
Build a message testing survey with 10 adaptable questions, clear answer options, a comprehension coding example, and a plan for comparing variants.
A message testing survey checks what people understand, believe and find relevant in a specific piece of communication. Show the message, ask for an unaided explanation, and then investigate clarity, credibility and relevance separately. Keep the underlying offer constant when the decision concerns wording.
Start with one communication decision
Imagine a fictional software service that turns uploaded meeting notes into a checklist. It does not attend meetings or assign work automatically. The team wants to compare two headlines while keeping the product description, price, layout and call to action the same.
Version A reads “Turn meeting notes into a project checklist.” Version B reads “Your meeting, ready for action.” The communication decision is whether readers correctly understand that they must supply the notes. Neither headline is a product the site is selling, and no consumer test of these examples has been conducted.
Before recruitment, write what result would change the headline, which misunderstanding would block adoption, and how you will code the primary outcome. The existing decision-brief guide covers that planning step in more detail. This page supplies the questionnaire that follows it.
10 questions to adapt
Ask only the questions needed for your decision. Consent and essential eligibility checks come before message exposure. Keep unaided questions before any wording that reveals how the product is supposed to work. The categories below are proposed templates, not a validated scale.
| ID and question | Response format | Decision supported |
|---|---|---|
| Q1. In your own words, what would this service do? | Open text; allow blank/unsure | Does the central function come through? |
| Q2. What would you need to do before the service could create a checklist? | Open text; allow unsure | Do people understand that notes must be supplied? |
| Q3. What, if anything, is unclear in this message? | Open text | What information or wording needs revision? |
| Q4. How clear or unclear is the message? | Very unclear; somewhat unclear; neither; somewhat clear; very clear; cannot judge | How clear does it feel, separately from accuracy? |
| Q5. How relevant is this service to a task you currently do? | Not at all; slightly; moderately; very; extremely relevant; cannot judge | Does the explained use case fit this audience? |
| Q6. How believable or unbelievable is the main claim? | Very unbelievable; somewhat unbelievable; neither; somewhat believable; very believable; cannot judge | Does the stated promise raise a credibility problem? |
| Q7. What most influenced your answer about believability? | Open text for every rating direction | Which evidence or wording caused doubt or confidence? |
| Q8. Which description best matches what you think the service does? | Turns supplied notes into a checklist; attends meetings and writes notes; assigns work automatically; none of these; cannot tell | Which specific aided interpretation is selected? |
| Q9. What would you need to know before deciding whether to try it? | Open text | What information is missing from the next page? |
| Q10. What, if anything, would make you decide against trying it? | Open text | Which barrier needs investigation? |
Use a coding rule that can detect the wrong promise
For Q1, a draft coding scheme might distinguish accurate, partial, incorrect and unsure. Accurate means the answer identifies the transformation of supplied notes into a checklist without adding attendance or automatic assignment. Partial captures some relevant function but leaves the input unclear. Incorrect includes one of the false capabilities.
For Q2, the key idea is that the user supplies meeting notes. Preserve the original answer and the code. If the primary analysis collapses codes into accurate versus other, report the fuller categories too so a favorable number cannot hide a specific misunderstanding.
Have reviewers code without seeing the headline condition where feasible. Resolve disagreements through a documented rule, and keep examples of difficult cases. AI can propose labels, but checking source text and adjudicating meaning remain necessary.
Keep exposure and analysis aligned
Random assignment to one message per person supports an independent-groups comparison when the study is implemented appropriately. Showing both headlines to everyone answers a different question and produces repeated or paired observations.
For an independent-group binary outcome such as accurate versus other, the message-comparison calculator can show a descriptive difference and its supported inference. It is not suitable for treating two repeated ratings from the same people as independent groups. It also cannot repair unequal exposure, an unsuitable sample or an outcome redefined after the results arrived.
Choose sample size through a power analysis for the intended comparison. The field-size planner helps with recruitment arithmetic once a usable base is justified; it does not choose that base scientifically.
Sources: Single versus repeated exposure
What to do when clarity and comprehension disagree
Suppose a future study finds that a message feels clear while many readers think the service attends meetings. That is not a reason to average the two outcomes into a reassuring total. The positive feeling may be attached to a false promise.
Likewise, correct understanding with low relevance can mean the communication worked and the audience does not need the offer. Rewriting the headline cannot be assumed to solve a product or audience mismatch.
Write the next action in terms of the actual finding: investigate the false capability, add the missing input requirement, refine the target audience, or plan a behavioral test. Avoid “winning message” language unless the predeclared decision rule and study quality support it.
How to adapt this template with AI
Give an assistant the factual offer, the decision and the exact versions. Ask it to preserve the offer facts, identify hypothesis-revealing questions and flag any item that measures two things. Do not ask it to invent customer responses to pick a winner.
Use the checker to prepare a review brief, then inspect the questions in the real survey software. Display logic, required fields, mobile layout and response order are part of the instrument. The Census standard discusses testing data-collection instruments; a clean draft alone is not the finished survey.
What is the difference from concept testing?
Concept testing investigates a proposed product, service or offer. Message testing investigates how an offer is communicated. If one version changes both the headline and the feature set, the result cannot isolate the effect of wording.
Can these questions predict conversions? They can reveal interpretation and stated judgments in the surveyed context. They do not establish a conversion rate, sales forecast or population-wide demand. Follow the reporting checklist to keep wording, dates, sample and limitations attached to any actual results.
Sources and scope
- Statistical Quality Standard A2U.S. Census Bureau
Primary standards for developing and pretesting data-collection instruments; not certification of these templates.
- Monadic versus sequential monadic survey designSurveyMonkey
Commercial research-platform explanation of single and repeated concept exposure. Supports the design distinction, not a universal sample-size rule.
Sources support the specific statements described above; they do not validate this publication’s rubric, guarantee a compliant execution, or replace context-specific professional advice.
What this page is: a research-methods guide, not a report of completed consumer fieldwork. Corrections and material revisions are recorded under the publication’s editorial standards.