Statistical Significance

Statistical significance is the line between a result and a coincidence, the discipline protecting tests from eagerness.

What statistical significance is

Statistical significance is the test’s answer to “or was that luck?”: a threshold saying the measured difference between variants is unlikely to be random noise, the line between a result and a coincidence.

Why statistical significance matters

Traffic split two ways never performs identically even when the variants are the same, so every test shows a difference; significance asks whether the difference means anything. Ignoring it turns testing into astrology, winners crowned by noise, then quietly failing at rollout, and the discipline is built precisely to protect the store from its own eagerness to see a result.

Reading tests honestly

  • Decide the finish line first: sample size and duration set before launch, not negotiated during
  • No peeking calls: checking daily and stopping at the first good number manufactures false winners
  • Full cycles: whole weeks, so weekend buyers and payday patterns are represented
  • Significant isn’t always meaningful: a real but tiny lift may not be worth the change

Frequently asked questions

What significance level should store tests use?

The conventional threshold is fine; the violations are what ruin tests, stopping early, running dozens of variants until one clears, or retesting losers until they win. Fixed rules honored beat exotic statistics adjusted mid-flight.

Can low-traffic stores reach significance?

Only on big effects: small conversion differences need traffic small stores don’t have, which is why their winning move is testing bold swings, offers, page structures, entire value propositions, on their highest-traffic pages, and accepting that button-color questions will simply never resolve at their scale.

Related terms

A/B Testing · Conversion Rate · Funnel Analysis