You have two ads in the same ad group. One is at 1.5% CTR, the other at 1.0%. Looks easy: the first one wins, you archive the second and move on.
That decision is a coin toss dressed up as analysis. And in ChatGPT Ads, run from OpenAI's Ads Manager, it hurts more than in Google Ads: volumes are small and every test takes weeks.
Why two CTRs aren't enough
A CTR is not a property of the ad: it is the outcome of a draw. With the same creative, two different days give different figures. You aren't comparing two numbers, but two noisy estimates.
And that noise has a known size: a count moves, by chance, by roughly its square root. An ad with 30 clicks could have had 24 or 36 with nothing having changed. That is why a 10-click gap between two small ads means nothing, while the same proportional gap with ten times the data does.
⚠️ Trap: the mistake isn't only picking the wrong ad. It's archiving an ad that was just as good, losing its testing slot and believing you learned something. False decisions contaminate the ones that follow.
What "significant" means, in plain words
When a difference is said to be statistically significant at p < 0.05, this is what is being said: if the two ads were actually identical, a gap as big as the one I'm seeing would turn up by pure chance less than once in twenty times. It isn't proof that one is better; it is a prudence threshold that rules out the most common coincidences.
Two nuances make it usable:
- Significant isn't important. With huge amounts of data, a 0.5% gap can be significant and still not be worth changing anything for. So a minimum effect size is required as well.
- Not significant isn't "they're the same". Almost always it means "I don't know yet". To claim two ads tie you need more sample than to declare a winner, not less.
The in-house method
This is the one we apply, and it is built for small volumes:
| Rule | Threshold | Why |
|---|---|---|
| Minimums per arm | 500 impressions, 20 clicks and 3 full days | Below that you don't even look: any figure is noise |
| Declare a winner | Significant (p < 0.05) and at least 10% in relative terms | The first rules out chance; the second, gaps not worth acting on |
| Declare a tie | Four times the minimums (2,000 impressions and 80 clicks) and no signal | Claiming there is no difference takes more data than finding one |
| Time limit | 28 days | An endless test blocks the slot for the next one |
| Conversion rate | Only decides with ≥ 5 conversions per arm | Below that it is the noisiest metric of all |
With enough conversions in both arms the conversion rate decides, because that is what you pay for; if they don't reach that floor, CTR decides, the signal that settles first.
💡 Ninja trick: the mental shortcut, valid when both arms have had similar impressions: the click gap has to beat twice the square root of the sum of the two counts. That is, in other words, the same calculation the statistical test performs.
A worked example
Two ads in the same group, with a similar share of impressions, at two moments and with the same CTR:
| Moment | Ad | Impressions | Clicks | CTR |
|---|---|---|---|---|
| Week 1 | A | 2,000 | 30 | 1.50% |
| Week 1 | B | 2,000 | 20 | 1.00% |
| Day 10 | A | 6,000 | 90 | 1.50% |
| Day 10 | B | 6,000 | 60 | 1.00% |
Week 1. A looks 50% better. Shortcut: 30 + 20 = 50 clicks; square root 7.1; twice that, 14.2. The actual gap is 10 clicks, below the bar. The formal test says the same: p ≈ 0.15, meaning something like that would come up by chance one time in seven. Verdict: not enough sample, nothing is touched.
Day 10. 90 + 60 = 150; square root 12.2; twice that, 24.5. The gap is 30 clicks and clears it. The formal test: p ≈ 0.014. The relative difference (50%) passes the required 10% comfortably and both arms are past the minimums. Verdict: A wins.
Same percentages, opposite decision. The only thing that changed was the amount of data, which is precisely what you cannot improvise.
The honest limit: this is an observational comparison
In ChatGPT Ads you cannot control how ads within a group rotate (checked by us, September 2026): the platform decides who sees each one. The two arms don't get identical audiences, so part of the gap may come from the split rather than from the creative.
It doesn't invalidate the test, but it forces you to be stricter: hence the high minimums, the three full days (so different moments of the week are included) and the extra requirement of a 10% relative gap. This is no laboratory experiment, and it's worth saying so.
What to do with each verdict
- Winner: it becomes the reference ad and the loser gets archived, never paused. An "activate all" revives anything paused and puts you back at square one.
- Tie: the older one stays, as it already has history, and the slot is freed for the next hypothesis.
- Not enough sample at 28 days: close it with no conclusion and write it down. It tells you that on your budget that variable can't be measured and that the next test should be bolder.
And one hygiene rule: one live test per ad group. With two at once you can't tell which one caused the change. How to build the variants is covered in "Ads and landing pages in ChatGPT Ads"; the weekly routine that chains them together, in "Optimising ChatGPT Ads without search terms".
📌 Do the maths before you start: on a small budget and with a click costing several euros (the figures going around come from agencies such as Choice OMG in 2026, not from OpenAI, and quote $3-5), reaching 20 clicks per arm can take weeks. Fewer variants at a time, and longer on each one.
What to take away
- A CTR is a noisy estimate: chance moves a count by roughly its square root.
- Minimums per arm before you look at anything: 500 impressions, 20 clicks and 3 full days.
- A winner only with p < 0.05 and a 10% relative gap; a tie only with four times the minimums; close at 28 days.
- Conversion rate decides only with 5 or more conversions per arm; otherwise CTR decides.
- With no control over rotation, the comparison is observational: that calls for more sample, not less.
- The loser gets archived, and only one live test per ad group.