Your winning ad might be a coin flip (check before you scale)

· has scaled a “winner” that turned out to be noise🪙 with an A/A coin-flip simulator, not a lecture

Published October 4, 2026

The 60-second version

  • Small tests produce big fake winners — two identical ads regularly differ by 10–25% on a week of modest traffic.
  • Size the test before launch: catching a 20% lift takes roughly 400 conversions per ad, not 40.
  • The fix: use real split tests, don’t peek or stop early, wait for conversion lag, and record the confidence range.
bottom line: size it first, read it once

Friday afternoon, the creative test results are in. Ad B — the new hook — converted at 2.20%. Ad A, the old one, at 1.80%. A 22% lift. The team pauses Ad A, moves the budget to B, and writes "new hook wins" in the learnings doc.

Four weeks later, B is converting at 2.0%, exactly where A always was. The hook didn't stop working. It never worked better in the first place. One week of data on 6,000 visitors per ad simply can't tell a 22% gap from luck.

Small tests produce big fake winners. heads, heads, heads… Every team running creative tests has crowned one, scaled it, and filed a "learning" that was just noise. The fix is not more tests. It's knowing how much data a test needs before anyone is allowed to call it.

The winner that wasn't

Creative test review
Ad A (old hook)108 / 6,0001.80% conversion rate, one week
Ad B (new hook)132 / 6,0002.20% conversion rate, one week
Apparent lift+22%What the learnings doc recorded
Confidence it's real88% (p = 0.12)Short of the usual 95% bar
Needed to confirm a 20% lift~19,600 visitors per ad≈ 392 conversions each — over 3x what they had
Next four weeksB 2.01% vs A 1.98%The gap was luck all along
88% sounds convincing. It means roughly a 1-in-8 chance of seeing a gap this big between two identical ads. A team testing every week will crown a fake winner like this roughly once every two months.

Two identical ads, one "winner"

The simulator below runs an A/A test: both ads have exactly the same 2% conversion rate. Any gap you see is pure chance. Press re-run a few times at the small traffic level, then switch to the large one.

LIVE · RE-RUN ITTwo identical ads, one week
Ad A45 conv / 2,800 visitors · 1.61%
Ad B56 conv / 2,800 visitors · 2.00%
Apparent lift
+24%
Confidence it's real
73%
AD B "WINS" — FAKE WINNER

A 24% gap between two identical ads. A dashboard would call this a winner; it is a coin landing heads a few extra times.

runs: 1 · fake 10%+ “winners”: 1
simulatedboth ads convert at exactly 2.0% — every gap you see is luck

At 400 visitors a day, gaps of 10–25% between identical ads show up all the time. At 4,000 a day, they shrink to a few percent. Nothing about the ads changed — only the amount of data. The smaller the test, the bigger the lies it can tell.


Why ad platforms crown winners too early

Uneven spend isn't a test

Put four ads in one ad set and the platform picks a favourite in the first day or two, then starves the rest. The loser never got enough traffic to lose fairly.

Peeking at the results

Checking daily and stopping the first day it looks 'significant' gives luck many chances to win. Commonly cited simulations put the fake-winner rate at 20–30% instead of 5%.

Too many variants

Test ten ads against a control and, by chance alone, one will usually look like a winner at the 90% level. More arms need more data — or a stricter bar.

Use the platform's real split-test tools for creative tests. Meta's A/B test feature and Google Ads experiments split people into separate groups, so each version gets a fair share of comparable traffic. A normal ad set or ad group is built to optimise, not to compare.


How much data a test needs

The amount of data depends on two things: how often people convert, and how small a difference you want to catch. Low conversion rates and small lifts need a lot of traffic. Try your numbers:

LIVE · DRAG ITHow big does your test need to be?
2%
20%

Relative lift: 20% means 2.0% → 2.4%

500
Visitors needed per ad
19,600
Conversions per ad
392
Days to read the test
39.2
TOO SMALL TO READ

This test won't give a trustworthy answer within a month. Test bolder ideas (bigger expected lifts), pool traffic into fewer variants, or test on a higher-volume metric like click-through rate.

live mathsrule of thumb for 95% confidence and 80% power: n ≈ 16 × p(1−p) ÷ (p × lift)²

The rule of thumb worth remembering: to catch a 20% lift, you need roughly 400 conversions per ad. Not 40. Not 100. That's why bold concepts are easier to test than small tweaks — a 50% lift needs only about 60 conversions per ad.

Wait for conversion lag. If customers often buy two or three days after clicking, the last days of a test are missing conversions. Read the test only after the lag window has closed, or late conversions will reshuffle the result.


Reading the result honestly

Once the test has reached its planned size, check significance once. This query computes the lift, confidence and a range for the true lift from a simple test table — the same maths as the significance tool below.

A/B result: lift, z-score and 95% range

A z_score above 1.96 (or below −1.96) clears the 95% bar. More useful still is the range: if diff_low_pp is below zero, the true effect could be nothing — or negative. For the "winner" above, z is 1.56 and the range runs from −0.10 to +0.90 points. That range includes zero, so the test proved nothing.


A test protocol that survives

Before, during and after a creative test

Process Flow
1

Write the question and the metric

One hypothesis, one primary metric — conversion rate, CPA or click-through rate. Decide before launch, not after you see the numbers.

2

Size it before you start

Use the calculator above: the lift worth catching sets the visitors and days needed. Too long? Test a bolder idea instead.

3

Split traffic properly

Use the platform's A/B test or experiment feature, with two to three variants at most.

4

Don't peek, don't stop early

Let it run to the planned size, plus the conversion-lag window. Daily checks are for broken links, not for winners.

5

Read once, record the range

Write the lift, the confidence and the range in the learnings doc. 'Not proven' is a valid result — it saves you from scaling luck.

Testing habits worth keeping pin these to the test tracker →

  • No winner below ~400 conversions per arm for a 20% lift — or below whatever your size check says. Smaller tests are 'directional', never 'proven'.
  • Test big swings, not button colours — new angles and formats produce lifts large enough to measure with normal budgets.
  • Re-check winners after scaling — compare the first four weeks at full budget with the test result. If the lift vanished, update the learnings doc.

Quick gut-check

One question. If you get it, the whole post clicks. 30 seconds, no stats degree

A new ad shows a 29% higher conversion rate after 3 days, on 45 conversions vs 35 from similar traffic. The team wants to scale it today. What do you say?


Frequently asked questions

How many conversions do I need for an ad A/B test?

It depends on the lift you want to detect. As a rule of thumb, catching a 20% relative lift needs roughly 400 conversions per ad; a 50% lift needs about 60; a 10% lift needs about 1,600. Smaller lifts need far more data.

Is 90% confidence good enough for ad tests?

It can be, for low-risk decisions you can reverse cheaply. Just know what it means: about a 1-in-10 chance of a gap that size between identical ads. If you run many tests, some of your "winners" will be luck. For decisions that move large budgets, use 95%.

Can I test creatives inside one ad set?

You can compare them, but it isn't a fair test: the platform shifts spend to its early favourite, so the others never get equal traffic. Use the platform's A/B test or experiment tools when you need a real answer.


The summary

  • Two identical ads regularly show 10–25% gaps on small samples — big lifts on small tests are usually luck.
  • Size every test before launch: low conversion rates and small lifts need a lot of traffic.
  • Catching a 20% lift takes roughly 400 conversions per ad.
  • Use proper split tests, don't peek, and wait for conversion lag before reading results.
  • Record the confidence range, not just the lift — "not proven" is a result.

Takeaways for your next report

  • Small tests produce big fake winners: identical ads often differ by 20% in a week.
  • Rule of thumb: ~400 conversions per ad to confirm a 20% lift at 95% confidence.
  • Stopping the first day a test looks significant can turn a 5% false-winner rate into 20–30%.
  • Ad sets optimise, they don't compare — use the platform's A/B test or experiment tools.
  • Test bold ideas: bigger expected lifts need far less traffic to prove.
stick this on your Monday report
Free tool

A/B Test Significance Playground

Enter visitors and conversions for control and variant to get a verdict, confidence, lift bounds and the sample size you actually needed.

Chinmay Raibagkar

Chinmay Raibagkar

About author →

Founder of DataLens AI. He helps non-technical teams read their ad and database numbers with confidence — which number to trust, what to do next, and what to ignore.

Glossary terms referenced