Changing the strategy of an important campaign "to see what happens" is expensive: it resets learning, mixes periods and leaves you with no counterfactual. Google Ads custom experiments solve all three: a copy of the campaign with the change, traffic split between them, the same period, and statistical significance. This lesson is the manual for using them on bidding questions.
Which questions deserve an experiment
- Target CPA versus target ROAS (with values already in place).
- A different target (−20%) to see what it costs you in volume.
- Its own strategy versus a portfolio.
- With and without broad match (module 1 of the Intermediate level).
- With and without AI Max (module 3 of this level; its own experiment).
- Maximise conversions versus target CPA in accounts with medium volume.
What does not deserve an experiment: small target changes (±10%), which are handled with the judge and its steps; or campaigns with no volume (the experiment will never reach significance).
Design
- One variable: only the strategy (or the target) changes. The same budget in both arms (the experiment splits it).
- A 50/50 split based on searches (each user always in the same arm).
- Duration: a minimum of 4 weeks or two complete conversion cycles or until you reach 100+ conversions per arm; whichever takes longest. With long latency, add a "cooling" week at the end that you do not judge.
- Timing: away from seasonal peaks and with no other changes planned for the campaign.
- A primary metric defined beforehand: the reference conversions and the business CPA/ROAS, not simply the ones the tag counts.
Reading the results
Google shows the metrics per arm with confidence intervals and marks when the difference is significant (usually at 95%). Read it in three steps:
- Is it significant? If not, there is no result — neither for nor against. Extend it or accept that the difference is smaller than the noise.
- How big? The effect size (CPA −12%, conversions +8%) and its interval. A significant effect of 2% does not justify a change that costs you a learning phase.
- In what? More conversions at the same CPA (volume), the same volume more cheaply (efficiency), or more volume more expensively (a business decision: do you want it?).
And the whole chain: an experiment that wins on tag conversions and loses on qualified leads (if you measure them) has not won.
Mistakes that invalidate an experiment
- Touching the original campaign or the test one during the experiment.
- An 80/20 split "to be careful": the small arm takes twice as long to gather data.
- Judging after a week.
- Changing two things (strategy + new ads).
- Running the experiment in peak season: the result does not repeat in the low season.
- A test arm with the "desired" target instead of a realistic one: you are comparing against the impossible.
Taking the winner into production
Google lets you apply the experiment (turn the winning arm into the campaign) or end it. If you apply it: the original campaign adopts the change and keeps its history; the test arm closes. There is a partial learning reset; plan for it (not at a peak) and keep an eye on it for two weeks with the judge in observation mode.
If the winner is "keep things as they are", document the result: a negative experiment saves you the same mistake next year (the memory idea from module 4 of the Intermediate level applies to strategies too).
💡 Ninja trick: the Smart Bidding judge (SBNS) can be left in TEST mode over both arms of an experiment: it records each arm's verdict every night without touching them, and when it ends you have, on top of Google's official result, the daily compliance series per arm — the way to see when one arm started to pull away from the other, not just whether it did.
What you should remember
- Experiment with the big questions (mode, a distant target, portfolios, broad match, AI Max); the small ones go to the judge.
- One variable, 50/50, 4+ weeks or two cycles or 100+ conversions per arm, away from peaks.
- Significance → size → nature of the effect; and the whole chain.
- Apply the winner away from a peak and watch it; document the negatives too.