Field note · causal inference

Avoid the trap of “fake” A/B tests in business

A randomized trial can answer a well-defined causal question. Before proposing one, check whether you can run it and what the result would actually tell you.

01

It’s a trial, not a split.

I prefer the full term, randomized controlled trial, because it reminds us what the design requires: random assignment, a control group, and a specific change to evaluate.

The question here is whether the exposure changed the conversion rate, allowing for chance variation between the groups.

An exposed group converting at 12 percent above an unexposed control group converting at 10 percent; the two-point gap is an estimate of the causal effect. exposed converts 12% not exposed converts 10%
Those 2 points estimate the causal effect. Random assignment makes the control group a useful comparison for what would have happened without the exposure. The estimate still has uncertainty.

02

Check whether the question can be tested.

When a decision is uncertain, asking for a test is understandable. The difficulty is that some decisions do not give you a workable treatment and control group.

“Can we just A/B test it?” — asked about a decision that cannot be A/B tested

03

Splitting traffic isn’t testing.

Sending half your users somewhere different is only part of the design. Decide what you will measure and how you will evaluate it before looking at the results.

An unreliable test

  • Split the traffic
  • Watch the dashboard daily
  • Pick the winning metric after looking
  • Stop when it looks good
  • Report only the win

A trial

  • Randomize users, not page views
  • Name one conversion event first
  • Set the MDE before you look
  • Stop at the sample size you planned
  • Apply the agreed decision rule

MDE is the minimum detectable effect: the smallest lift worth shipping, say +1.5 points on a 5% baseline. Choose it first. It sets your sample size, and it is what separates real from worth it: an effect can be statistically significant and still too small to justify the work.

04

Wall 1

There is no multiverse.

You cannot randomize the entire decision to launch a new product into two versions of the same market. You may be able to test parts of the launch, but those tests answer narrower questions.

A timeline reaching a launch point, from which one solid branch is the world that happened and three faded branches are worlds you can never observe. worlds you can’t see launch the world you got
This launch gives you one observed outcome. A comparison needs to be designed separately.

05

Wall 2

The arithmetic says no.

Purchase conversion may be the outcome you care about, but a low purchase rate can make the required sample impractical.

Take 2,000 visitors a week and a 2% purchase rate. Detecting a 10% lift (an MDE of 0.2 points) needs about 78,400 users per group.

Users needed per group, for a 10% lift

Bought something 78,400

2% baseline · MDE 0.2pp · about 18 months

Reached checkout 6,400

20% baseline · MDE 2pp · about 6 weeks

n ≈ 16 × p(1−p) / MDE², at 95% significance and 80% power

Eighteen months. The feature gets rewritten twice, and whoever asked for the test has changed teams.

06

Test one link, not the whole chain.

So narrow the question. Make the conversion event an earlier step in the journey, one with the traffic to support a real test. That is the second bar: six weeks instead of eighteen months.

A user journey

  1. Lands on the page
  2. Sees the new offer
  3. Reaches checkout
  4. Buys
  5. Comes back next month
The trial settles the highlighted link. The rest of the chain is not in the test and never was.

This test estimates the effect on checkout rate. It does not establish the effect on purchases or repeat business; those need additional evidence.

State that limit alongside the result. Otherwise, an improvement in checkout rate can easily be repeated as evidence of higher revenue.

07

No trial possible? You still have tools.

Depending on the setting, these designs may help. Some can use randomization; others rely on assumptions about the comparison group or what would have happened over time.

  • Staged rollout Introduce the change in some regions first, with others available for comparison.
  • Switchback test Alternate the change on and off over time, accounting for effects that may carry over between periods.
  • Difference-in-differences Compare the change in your group against the change in a similar untouched group.
  • Regression discontinuity Some rule has a hard cutoff. Compare people just above it to people just below.
  • Synthetic control Build a weighted twin of the treated market out of the untreated ones.
  • Pre/post analysis Compare before and after, while checking for other changes that could explain the result.

Discuss the evaluation before launch. You may need to preserve a comparison group, schedule a rollout, or collect a baseline that cannot be reconstructed later.

The job

Say what the test can answer.

If the test cannot support the claim, say so before the result goes into the presentation. Explain what you did learn and what remains uncertain.