1. Write one falsifiable hypothesis
Name the audience, change, expected behavior, primary metric, and reason. For example: “For eligible mobile visitors, showing the expected number of questions before the start will increase quiz starts per eligible view because the effort becomes clearer.”
Do not write “Variant B will convert better.” It does not identify which behavior should change or why. A precise hypothesis also helps the team choose a metric close to the changed experience instead of defaulting to total leads.
2. Change one decision-relevant variable
Keep the current experience as the control and change one interpretable element in the variant. Suitable tests include the promise before the quiz, the explanation of progress, the first-question wording, the position of an optional contact gate, or the next step on one result.
Changing the headline, question count, visual design, result copy, and contact form together may produce a different outcome, but it cannot show which change mattered. If the whole journey must be redesigned, treat it as a broader replacement test and describe the conclusion at that level.
- Do not remove necessary privacy or accessibility information to create a shorter variant
- Do not give one variant a different traffic source or campaign promise
- Keep scoring and result routes stable unless they are the stated test variable
- Verify that both variants load and record events before sending production traffic
3. Randomize assignment and keep it stable
Randomly assign eligible experimental units to the control or variant. For most quiz tests, the unit is a visitor or account, not an individual page view. Store the assignment so returning participants do not switch variants halfway through the journey.
Run both versions concurrently. A test that shows A this week and B next week is exposed to campaign changes, weekday patterns, news, seasonality, and other time effects. Randomization does not remove every source of bias, but it distributes ordinary variation more defensibly than sequential exposure.
4. Predefine one primary metric and guardrails
Choose the primary metric that directly answers the hypothesis. Add guardrails that can stop a superficially positive result from harming the rest of the journey.
| Test focus | Possible primary metric | Useful guardrails |
|---|---|---|
| Pre-quiz promise | Starts divided by eligible views | Completion, result delivery, qualified yield |
| Question wording | Progress past the changed question | Result distribution, error rate, accessibility |
| Contact-gate position | Intentional contact submissions | Result views, consent quality, qualified outcome |
| Result-page next step | Next-step actions per result view | Lead quality, opt-out, support complaints |
| Scoring or routing | Correctly qualified outcomes | Result balance, manual overrides, false exclusions |
Related: Check the scoring rules before testing a routing change
5. Calculate feasibility before launch
Determine the baseline rate, smallest effect worth acting on, acceptable false-positive risk, desired power, allocation, and expected eligible traffic before choosing a sample size. Use an analysis method suitable for the metric and experiment design, with statistical review when the decision is consequential.
If the required sample cannot be reached in a useful period, do not weaken the decision rule after seeing early results. Test a larger, more meaningful change, combine the experiment with structured usability research, or record that the site does not currently have enough traffic for this question.
6. Run the test without moving the goalposts
Before exposure, record the hypothesis, variants, eligibility rule, assignment unit, metrics, sample plan, start condition, stop condition, exclusions, and analysis. Monitor data loss and severe harm, but do not repeatedly declare a winner whenever the chart crosses a preferred threshold unless the method explicitly supports sequential monitoring.
Record campaign launches, outages, tracking changes, and material audience shifts. If one event fails only in a variant, repair the instrumentation and decide whether the contaminated period must be excluded under the rule written before analysis.
7. Interpret the full journey and retain the learning
Report the effect estimate and uncertainty, not only the winning label. Check the predeclared guardrails and relevant result or device segments. Treat unexpected subgroup patterns as hypotheses for later work unless the test was designed and powered to answer them.
A higher completion rate is not automatically a business win. If the variant reduces result usefulness, consent quality, or downstream qualification, the responsible decision may be to keep the control or redesign the test. Store the brief, QA record, result, decision, and follow-up question so future tests do not repeat the same uncertainty.
Evidence
How to reproduce the method
The experiment is reproducible from a dated brief containing the hypothesis, control, changed variable, eligibility rule, random assignment unit, event definitions, primary metric, guardrails, feasibility assumptions, stop rule, QA results, exclusions, analysis, and final decision. Another analyst can inspect those artifacts without relying on a dashboard screenshot alone.
Limitation
Where this conclusion stops
This guide does not prescribe a universal sample size, test duration, significance threshold, or statistical model. Low traffic, repeated visitors, cross-device identity, network effects, multiple comparisons, delayed outcomes, and changing campaigns may require a different design or specialist statistical review.
Sources and verification
What this guide relies on
- The Lead Quiz Review editorial methodology
- NIST, Engineering Statistics Handbook
- NIST, Choosing an experimental design
- NIST, Completely randomized designs
Sources and method checked August 11, 2026. External standards are linked to their primary publishers.