A/B testing tools or an A/B testing programme: what you lack

Most brands that say they need a testing tool already have one. What they lack is the hypothesis pipeline that feeds it and the discipline that reads it. A tool splits traffic and calculates significance. It does not decide what to test, and deciding what to test is where nearly all the value sits.

What the tool does

Splits traffic, serves variants, records outcomes, calculates confidence. Genuinely valuable and genuinely commoditised. The differences between the major tools matter far less than any vendor comparison suggests, once you are past the basic question of whether it can do server-side or client-side rendering for your stack.

The decision that actually matters is whether it assigns at the session level or the customer level, because that governs whether you can safely test pricing. Everything else is preference.

What the tool does not do

Generate hypotheses. A tool has no opinion about your funnel. Hypotheses come from analytics review, session recordings, survey coding and support tickets, and building that pipeline is a weekly habit rather than a purchase.

Prioritise. Ranking by expected value means traffic on the page multiplied by expected effect size multiplied by confidence in the evidence. No tool computes that for you, and getting it wrong is how a year gets spent on tests that could never have mattered.

Stop you peeking. Every tool shows a live result. Only a process stops someone calling it on day four.

Read the right metric. A tool reports what you configured. If you configured conversion rate, it will happily tell you a discount test won.

The gap in numbers

The base rate explains why prioritisation outranks execution. Optimizely’s analysis of more than 127,000 experiments puts the average win rate near 12%, and ConversionTeam’s audit of 2,288 tests found 19.1% reached statistical significance per test.

If four in five tests will be inconclusive regardless of how well they are built, then the only lever that meaningfully changes your output is choosing better tests. That is a research and judgement problem, and no software solves it.

 

Tool

Programme

Traffic splitting and stats

Uses the tool

Hypothesis generation

 

Prioritisation by expected value

 

Test build and cross-device QA

 

Discipline on runtime and sample

 

Honest reading, including losses

 

Roadmap re-ranking on results

 

Monthly cost

Lower

Higher

The tell that a tool is not your problem

Look at your test log for the last twelve months.

If it is empty or nearly so, the constraint is not tooling. It is that nobody owns the roadmap, or nobody has authority to approve a change, or there is no research producing hypotheses. Buying a better tool changes none of those.

If it is full of small tests with inconclusive results, the constraint is prioritisation. You are testing things too small to detect, and a better tool measures them no more successfully.

If it is full of winners with no revenue impact, the constraint is the metric. You are reading conversion rate where you should be reading revenue per visitor.

Only if it is full of well-prioritised tests hitting genuine technical limits is the tool actually the problem, and that situation is rare.

Where tool choice does genuinely matter

Two cases, and only two.

Pricing. Assignment must persist at customer level across sessions and devices, or one shopper eventually sees two prices. Most general-purpose tools do not do this, and discovering it afterwards is expensive.

Server-side rendering. If your stack renders server-side, a client-side tool introduces flicker that biases results toward the control. That is a technical constraint rather than a preference.

Outside those two, tool choice is close to irrelevant next to whether anyone is generating good hypotheses. ConversionTeam’s audit of 2,288 tests found comparable base rates regardless of platform.

What velocity is worth once the pipeline exists

Research from Harvard Business School cited in industry analysis found companies adopting systematic A/B testing see performance improvements of 30% to 100% within a year, with the highest testing velocity compounding fastest.

Velocity, not tooling. The brands compounding fastest are running more well-chosen tests, not running them in better software.

Keep the tool you have. Build the pipeline. You can read the full testing scope for what that pipeline contains.

 

Enjoyed this article? Stay informed by joining our newsletter!

Comments

You must be logged in to post a comment.

About Author