Guide library

Original tool test · September 23, 2026

What our survey question checker catches—and what it misses

We ran twelve selected questions through our actual wording rules. Inspect the inputs, alerts, missed problems and downloadable results.

Tool testingQuestion wordingReproducible examples

A wording alert can help you spot a problem. It cannot tell you whether a question measures what you intended. To show the difference, we ran twelve deliberately chosen inputs through version 1.0 of the same review function used by our Survey Question Checker. Some examples triggered useful warnings; others exposed gaps or needed context. The complete inputs and recorded outputs are available below.

What we ran and what counts as evidence

We selected twelve English question blocks to exercise leading assumptions, evaluative wording, time periods, combined judgments, numeric ranges and personal-data prompts. The cases were chosen for teaching value. They were not sampled from customer questionnaires and were not labeled by a panel of independent raters.

On September 23, 2026, each block was passed separately to the actual reviewSurvey function, rules version 1.0. The test did not change the checker's rules. We captured the parsed questions, rule codes, matched wording and error output. All twelve inputs parsed without input errors. This is a source-function test; it is not a claim that every browser, screen reader or device was tested.

The CSV separates the observed alerts from our interpretation. The alerts are repeatable software outputs. The explanation of why a question needs attention is editorial analysis. Neither is a finding about how real respondents would answer.

Useful alerts: praise, time periods and overlapping ranges

“How much do you love our amazing checkout?” returned leading-assumption and evaluative-language alerts. The wording presumes affection and supplies praise. A starting revision is “What is your overall impression of the checkout?” Choose answer options that match the decision and allow unfavorable or uncertain reactions. That revision still needs a pretest.

“How often do you visit our website?” returned a time-period alert. “How often did you visit our website in the last seven days?” returned none. Adding a time window makes the intended recall period explicit, but seven days is only an example. A rare behavior may need a different window, and a longer window may be harder to remember accurately.

“How much did you spend last month?” with options “0–50; 50–100; Cannot judge” returned an overlap alert. If endpoints are inclusive, exactly 50 fits two choices. Specify non-overlapping boundaries and a currency. The missing currency did not produce its own alert, so fixing the detected issue would not complete the review.

Selected inputs and actual alerts from rules version 1.0
CaseQuestion or changeObserved signals
C01How much do you love our amazing checkout?Leading assumption; evaluative wording
C02How often do you visit our website?Define the time period
C03Add ‘in the last seven days’No listed patterns flagged
C04How satisfied are you with speed and reliability?Check the other direction; possible combined judgment
C06Spending options: 0–50; 50–100; Cannot judgePossible overlapping numeric options
C12What is your email address?Review the need for personal data

When flagged wording is the thing you need to test

Case C07 asked “How believable is the phrase "excellent value"?” The checker flagged “excellent” as evaluative wording. In this question, however, the quoted phrase is the stimulus. Removing that adjective would change what the respondent is evaluating.

Keep the intended stimulus intact and review the question around it. Decide whether you need a believability rating, an unaided explanation, or both. A possible wording is “How believable or unbelievable is this claim?” with the exact claim shown separately. That is a design choice, not an automatic correction.

This case illustrates why we do not show a pass mark or a bias score. Counting alerts would penalize some intentional wording while giving a clean-looking result to some problematic questions.

Three problems that produced no alert

Case C08, “How much money did you save by switching?”, produced no alerts. It assumes a saving occurred. First establish whether spending fell, stayed the same, rose, or cannot be compared. Only ask for an amount where that follow-up makes sense.

Case C09, “Would you rate our service highly because it is trusted by experts?”, also produced no alerts. It supplies an endorsement as a reason for a favorable answer. If the objective is unaided evaluation, remove that cue. If the endorsement itself is the stimulus, hold its presentation constant and measure its interpretation separately. The checker cannot establish whether the endorsement is true.

Case C10, “How would you rate the price, delivery speed, customer support?”, produced no alerts despite combining three attributes. This version's combined-judgment rule recognizes some conjunction patterns, not every comma-separated list. Ask separate questions if different answers would lead to different decisions.

No alert is not the same as no problem
CaseObserved outputManual review still needed
C08No listed patterns flaggedDoes the respondent actually have savings to report?
C09No listed patterns flaggedIs a favorable cue changing the intended task?
C10No listed patterns flaggedCould price, speed and support receive different ratings?

What the remaining cases tell us

C05 named both easy and difficult, supplied balanced options, and included “Did not try.” It returned no alerts. C11 used the established phrase “research and development” and returned none. These observations show two accepted input forms; they do not validate either item for a particular study.

C12 asked for an email address and triggered the personal-data reminder. A reminder cannot decide whether collection is necessary or handle consent for you. Keep respondent identifiers out of prompts sent to an outside AI service. This test used a question about an address, never a real person's address.

Repeat the check, then review the whole instrument

Open the downloaded CSV in a spreadsheet. Copy one input cell into the checker, preserving the line break before any “Options:” line. Choose Check wording and compare the displayed labels with the observed_signals column. The rule_codes column gives the stable identifiers from the recorded source output. Use rules version 1.0 for a like-for-like comparison; later versions may behave differently.

Then review the questionnaire without relying on those alerts. Ask what each question assumes, whether it requests one judgment, which answers are possible, and whether earlier questions teach the answer. Keep the original and proposed wording side by side, with the reason for each change. Pretest understanding with relevant people before fieldwork.

Report a discrepancy through our editorial contact with the case ID, rules version and non-sensitive input. A new rule or a corrected observation should receive a new version and a dated change note. We do not calculate a sensitivity, precision or accuracy percentage from these hand-picked cases.

Sources and scope

  1. Test inputs and recorded outputs, version 1.0Message Test Bench

    Our original twelve-case software run, September 23, 2026. Supports the output descriptions, not population claims or tool validation.

  2. Writing Survey QuestionsPew Research Center

    Question-writing and response-option guidance. It does not test, endorse or validate our checker.

  3. Statistical Quality Standard A2U.S. Census Bureau

    Instrument-development and pretesting standard; informs the distinction between a mechanical wording check and questionnaire testing.

What this page is: an original software test note using deliberately constructed questions. The outputs were recorded from rules version 1.0. No consumers were recruited and no questionnaire was validated. AI assisted the test preparation and explanation; this is not an independent expert review. Corrections and material revisions are recorded under the publication’s editorial standards.