When testing pays, and when it is theatre
A/B testing is worth it when three conditions hold: enough traffic to reach significance in weeks, a change big enough to matter, and the discipline to accept the result. Miss any one and testing becomes theatre. The standard I hold reports to is scientific rigour and learnings, and the sharpest example I have run was a personalised email play whose conversion difference was not statistically significant, a result worth more than a fake win, because it stopped us scaling a play that did not work.
The screening question I use on other people's claims applies to your own dashboard too. A candidate once told me they lifted campaign ROI from 3:1 to 4.5:1 by changing a product image, but it was not A/B tested and had no clear hypothesis, which makes it a story, not a result. Untested lifts have a hundred fathers: seasonality, mix shift, price changes. If you would not accept the claim from a stranger, do not accept it from your own memory.
The same skepticism, applied to hiring, is half of screening growth candidates: operators volunteer the caveats, storytellers volunteer the lift.Build the infrastructure before the crisis
The lesson I paid most for: after a client funnel's conversion rate halved overnight, I spent roughly 72 hours across a weekend retrofitting A/B tests just to isolate what had changed, barely sleeping. Testing infrastructure you build during a crisis costs ten times what it costs in peacetime. Wire the harness early, even if you run few tests, because its second job is diagnosis when something breaks.
In practice my defaults are pragmatic: a clear hypothesis or no test, decisive splits like 20/80 run hard for a defined window when a decision is urgent, and judgment instead of testing below the traffic floor, per the CRO page. Tests come from a ranked queue, biggest suspected leak first, which is what the audit produces and how the wider experiment cadence stays honest. Most of what people burn tests on, button colours on landing pages, is better solved by the structural rules on the landing page page, and testing reserved for the questions structure cannot answer.
Finally, keep a written log of every test, including the nulls. Institutional memory is the compounding asset here: the third growth lead to inherit a funnel without a test log re-runs the same three losing experiments the first two already paid for, which is the quiet way strategy resets to zero every eighteen months.
FAQ
When is A/B testing not worth it?
Below the traffic needed for significance within weeks, for trivial changes, or when the team will overrule losing results anyway. In those cases fix obvious leaks by judgment and save testing for genuinely uncertain, high-stakes questions.
What makes a good A/B test?
A written hypothesis, a change big enough to detect, a pre-committed run window, and acceptance of the result, including the null. A non-significant result that stops a bad play from scaling is a win, not a failure.
How long should an A/B test run?
Pre-commit the window before starting: long enough for significance at your traffic, covering full weekly cycles. For urgent decisions a hard split run over a defined short window beats an open-ended test nobody ever calls.