This site only uses technical cookies required for it to work: no tracking, no profiling. Cookie Policy

Skip to content
All terms

What is statistical significance?

It shows how incompatible a result is with the null hypothesis, not that the effect is true, large or worth deciding on.

Statistical significance is the measure of how incompatible an observed result in the data is with a reference statistical model, usually the null hypothesis that denies the effect being tested: if that incompatibility crosses a conventional threshold, most often a p-value below 0,05, the result is declared "statistically significant". That threshold is a convention inherited from early twentieth century statistics, not a law of nature, and it should be treated as such. The point a decision maker has to hold onto is what the number does NOT say: the p-value is not the probability that the hypothesis under study is true, nor the probability that the result is due to chance alone. It only measures how surprising the data would be if the null hypothesis were true. Confusing it with proof of truth, or with the size of the effect, is the costliest mistake a committee can make before signing off on a spend based on that number.

The threshold does not replace judgment

In 2016 the American Statistical Association published an official statement, authored by Ronald L. Wasserstein and Nicole A. Lazar, to correct years of misuse of the p-value in applied research. The principle most relevant to business is the third one: scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. A result can be significant and economically irrelevant: with a large enough sample, as often happens in the conversion tests described under A/B testing, even a 0,1% difference comes out statistically significant without justifying any implementation spend. The ASA also notes that, by itself, a p-value does not provide a good measure of evidence: that is why the confidence interval, which shows a range of plausible effect sizes, is almost always the more useful figure to bring to an investment committee. There is also the multiple comparisons problem: testing many metrics at once, some will come out significant by chance alone, and a multiple-comparisons correction only holds if the family of tests is declared up front, because it does not repair a hypothesis picked after seeing the data.

An enterprise example

An Italian scaleup records a significant sales increase in the region where it launched a new campaign, with a p-value below 0,01. The sales committee wants to roll the campaign out everywhere right away, but in that same window a local competitor closed two physical stores in that region. The correlation between campaign and sales is real and statistically robust, but it does not establish on its own how much of the effect comes from the campaign: the competing cause, the store closures, could explain most of the increase. Only by isolating the two causes, through a proper experimental design or the techniques described under Causal AI, can the company establish how much of the result is actually attributable to the campaign before authorizing a nationwide rollout.

Why it matters for decision makers

Whoever signs off on a decision based on a number must answer three questions, not one: is the result statistically significant, is it large enough to justify the spend, and does it have a plausible causal explanation beyond the observed correlation. A low p-value only answers the first. The other two remain a human judgment that no numerical threshold can delegate away, and they are also where a poorly chosen metric, as discussed under Goodhart's law, produces results that are significant and irrelevant at the same time. Treating "significant" as a synonym for "true, large and causal" is the mistake the ASA put in writing precisely because, in business practice, it is the costliest one.

  • A/B test · A randomized controlled experiment applied to a decision, not a CRO guide: when it applies and when it does not.
  • Causal AI · AI that models cause and effect, not just correlation: it answers what would happen IF you acted, not just what is associated with what.
  • Goodhart's law · When a metric becomes the target, it stops measuring: what Goodhart's law implies when you accept an AI project.
  • "If you can't measure it, you can't manage it" · The maxim circulates cut in half: Deming cited it only to call it a costly myth, never to endorse it.
  • Survivorship bias · Drawing conclusions only from the cases that made it to you, while ignoring the ones that did not.

A term that hits close to home? Let's talk.

CONTACT ME