Original tool test · September 23, 2026
What our survey question checker catches—and what it misses
We ran twelve selected questions through our actual wording rules. Inspect the inputs, alerts, missed problems and downloadable results.
A wording alert can help you spot a problem. It cannot tell you whether a question measures what you intended. To show the difference, we ran twelve deliberately chosen inputs through version 1.0 of the same review function used by our Survey Question Checker. Some examples triggered useful warnings; others exposed gaps or needed context. The complete inputs and recorded outputs are available below.
What we ran and what counts as evidence
We selected twelve English question blocks to exercise leading assumptions, evaluative wording, time periods, combined judgments, numeric ranges and personal-data prompts. The cases were chosen for teaching value. They were not sampled from customer questionnaires and were not labeled by a panel of independent raters.
On September 23, 2026, each block was passed separately to the actual reviewSurvey function, rules version 1.0. The test did not change the checker's rules. We captured the parsed questions, rule codes, matched wording and error output. All twelve inputs parsed without input errors. This is a source-function test; it is not a claim that every browser, screen reader or device was tested.
The CSV separates the observed alerts from our interpretation. The alerts are repeatable software outputs. The explanation of why a question needs attention is editorial analysis. Neither is a finding about how real respondents would answer.
Useful alerts: praise, time periods and overlapping ranges
“How much do you love our amazing checkout?” returned leading-assumption and evaluative-language alerts. The wording presumes affection and supplies praise. A starting revision is “What is your overall impression of the checkout?” Choose answer options that match the decision and allow unfavorable or uncertain reactions. That revision still needs a pretest.
“How often do you visit our website?” returned a time-period alert. “How often did you visit our website in the last seven days?” returned none. Adding a time window makes the intended recall period explicit, but seven days is only an example. A rare behavior may need a different window, and a longer window may be harder to remember accurately.
“How much did you spend last month?” with options “0–50; 50–100; Cannot judge” returned an overlap alert. If endpoints are inclusive, exactly 50 fits two choices. Specify non-overlapping boundaries and a currency. The missing currency did not produce its own alert, so fixing the detected issue would not complete the review.
| Case | Question or change | Observed signals |
|---|---|---|
| C01 | How much do you love our amazing checkout? | Leading assumption; evaluative wording |
| C02 | How often do you visit our website? | Define the time period |
| C03 | Add ‘in the last seven days’ | No listed patterns flagged |
| C04 | How satisfied are you with speed and reliability? | Check the other direction; possible combined judgment |
| C06 | Spending options: 0–50; 50–100; Cannot judge | Possible overlapping numeric options |
| C12 | What is your email address? | Review the need for personal data |
When flagged wording is the thing you need to test
Case C07 asked “How believable is the phrase "excellent value"?” The checker flagged “excellent” as evaluative wording. In this question, however, the quoted phrase is the stimulus. Removing that adjective would change what the respondent is evaluating.
Keep the intended stimulus intact and review the question around it. Decide whether you need a believability rating, an unaided explanation, or both. A possible wording is “How believable or unbelievable is this claim?” with the exact claim shown separately. That is a design choice, not an automatic correction.
This case illustrates why we do not show a pass mark or a bias score. Counting alerts would penalize some intentional wording while giving a clean-looking result to some problematic questions.
Three problems that produced no alert
Case C08, “How much money did you save by switching?”, produced no alerts. It assumes a saving occurred. First establish whether spending fell, stayed the same, rose, or cannot be compared. Only ask for an amount where that follow-up makes sense.
Case C09, “Would you rate our service highly because it is trusted by experts?”, also produced no alerts. It supplies an endorsement as a reason for a favorable answer. If the objective is unaided evaluation, remove that cue. If the endorsement itself is the stimulus, hold its presentation constant and measure its interpretation separately. The checker cannot establish whether the endorsement is true.
Case C10, “How would you rate the price, delivery speed, customer support?”, produced no alerts despite combining three attributes. This version's combined-judgment rule recognizes some conjunction patterns, not every comma-separated list. Ask separate questions if different answers would lead to different decisions.
| Case | Observed output | Manual review still needed |
|---|---|---|
| C08 | No listed patterns flagged | Does the respondent actually have savings to report? |
| C09 | No listed patterns flagged | Is a favorable cue changing the intended task? |
| C10 | No listed patterns flagged | Could price, speed and support receive different ratings? |
What the remaining cases tell us
C05 named both easy and difficult, supplied balanced options, and included “Did not try.” It returned no alerts. C11 used the established phrase “research and development” and returned none. These observations show two accepted input forms; they do not validate either item for a particular study.
C12 asked for an email address and triggered the personal-data reminder. A reminder cannot decide whether collection is necessary or handle consent for you. Keep respondent identifiers out of prompts sent to an outside AI service. This test used a question about an address, never a real person's address.
Repeat the check, then review the whole instrument
Open the downloaded CSV in a spreadsheet. Copy one input cell into the checker, preserving the line break before any “Options:” line. Choose Check wording and compare the displayed labels with the observed_signals column. The rule_codes column gives the stable identifiers from the recorded source output. Use rules version 1.0 for a like-for-like comparison; later versions may behave differently.
Then review the questionnaire without relying on those alerts. Ask what each question assumes, whether it requests one judgment, which answers are possible, and whether earlier questions teach the answer. Keep the original and proposed wording side by side, with the reason for each change. Pretest understanding with relevant people before fieldwork.
Report a discrepancy through our editorial contact with the case ID, rules version and non-sensitive input. A new rule or a corrected observation should receive a new version and a dated change note. We do not calculate a sensitivity, precision or accuracy percentage from these hand-picked cases.
Sources and scope
- Test inputs and recorded outputs, version 1.0Message Test Bench
Our original twelve-case software run, September 23, 2026. Supports the output descriptions, not population claims or tool validation.
- Writing Survey QuestionsPew Research Center
Question-writing and response-option guidance. It does not test, endorse or validate our checker.
- Statistical Quality Standard A2U.S. Census Bureau
Instrument-development and pretesting standard; informs the distinction between a mechanical wording check and questionnaire testing.
What this page is: an original software test note using deliberately constructed questions. The outputs were recorded from rules version 1.0. No consumers were recruited and no questionnaire was validated. AI assisted the test preparation and explanation; this is not an independent expert review. Corrections and material revisions are recorded under the publication’s editorial standards.