The manual way comes first, in Google Ads itself. Then the same problem as Goldbeater finds it, every day, without anyone asking.
Google Ads A/B testing splits one campaign and compares the halves.
An experiment runs a change beside the campaign it would change. Google splits the campaign’s traffic or budget between the original and a version with your change in it, both run over the same weeks, and you compare them at the end. If the change did better, you apply it to the original campaign or run the new version in its place.
That is an A/B test, and it answers what a before-and-after can’t. A season, a sale or a new competitor moves both halves at once, so what’s left between them is the change.
Experiments are set up under Campaigns › Experiments, a page Google used to call Drafts and experiments. All experiments lists every one in the account. Google’s page names seven kinds today:
| Kind | What it tests | Campaigns |
|---|---|---|
| Custom experiment | Smart Bidding, keyword match types, landing pages, audiences and ad groups, in a trial copy of the campaign | Search, Display, Video |
| Ad variation | One change of ad text or URL across many ads, such as one phrase swapped for another | Search |
| AI Max experiment | AI Max’s search term matching and asset optimization, on half of one campaign, with no copy made | Search |
| Performance Max experiment | What Performance Max adds beside your other campaigns, a move to it from Standard Shopping, or a setting or asset inside it | Performance Max |
| Video experiment | Which of your video ads is more effective on YouTube | Video |
| Demand Gen experiment | Which images or videos draw the most interest | Demand Gen |
| App asset experiment | The text and images inside an App campaign | App |
Two of them cover most of what a Search advertiser tests. A custom experiment tests a change to the campaign itself: how it bids, what it matches, where it lands. An ad variation tests a change to ad text, and Google recommends it over a custom experiment for Search ads. Google’s pages on custom experiments differ on Demand Gen and Hotel Ads campaigns, and agree that an App or a Shopping campaign can’t take one. Performance Max experiments have a set of their own: uplift, upgrade and optimization.
How to set up a Google Ads campaign experiment
First check the campaign can take a custom experiment. Google lists what it needs:
- Its own budget and bid strategy. The campaign has to be enabled, on an individual daily budget and a standard bid strategy. One on a shared budget or a portfolio bid strategy may not appear in the list.
- No expanded text ads. Google can’t create an experiment for a campaign that holds expanded text ads or older text ads, in paused and removed ad groups too. Remove them and allow 24 to 48 hours. The expanded text ads guide shows how to find them.
- No other experiment on those dates. A campaign can have five scheduled, and only one running at a time.
- Go to Campaigns › Experiments, select the plus button, then Custom.
- Select Custom experiment and the campaign type, then Continue.
- Name the experiment, choose the campaign to test, and give its copy a suffix. Google creates the trial campaign from the original.
- Make the change you’re testing in the trial campaign, and choose up to two goals to measure it by. Google’s example is Clicks and Increase.
- Under Experiment split, set the share of traffic and budget the trial takes. Google recommends 50%, for the best comparison between the two.
- Under Advanced, choose how people are assigned: Cookie-based or Search-based.
- Under Schedule the experiment, set the start and the duration, then select Save.
There’s a shorter way in: edit an enabled campaign’s settings and select Save as experiment, which builds the trial from the change you just made.
- Cookie-based assigns each person to one side, so nobody sees both. Google recommends it. If the campaign uses audience lists, Google says each should hold at least 10,000 people, or the results may be less accurate.
- Search-based assigns each search, so one person can meet both sides on different searches. Google says it may reach a statistically significant result faster.
- A Display campaign always splits by cookie.
Set it to run at least four weeks. Google recommends 4 to 6 weeks where a result can’t yet be determined, and says to allow 7 to 14 days for the trial side to stabilize. A trial that stops serving may be in its bid strategy’s learning phase, typically 7 days, which the learning phase guide covers. The first end date has to fall 2 to 12 weeks after the start, and you can extend it while the experiment runs. Once it ends, it can’t be resumed.
Then leave both sides alone. The split can’t be changed once the experiment is set up, changes to the original don’t reach the trial, and Google says changes to either side make the result harder to read. It also advises against running several experiments at once, since they can interfere with each other.
Google Ads ad variations test one change of ad text across campaigns
An ad variation changes the same words in many ads at once and shows the changed ads to a share of people. Google’s example changes a call to action from “Buy now” to “Buy today”. It works on responsive search ads.
- Go to Campaigns › Experiments, open the Ad variations tab and select the plus button.
- Under Select ads, choose All campaigns or one campaign, and under Target ad type, Responsive search ads. Filter ads narrows it by headline, description, display path or final URL.
- Under Create variation, choose Find and replace, Update text or Update URLs. Variations are case sensitive.
- Name it, set its start and end dates, and under Experiment split enter the share of the campaign’s budget it takes. Then select Create variation.
People are split by cookie, so a person sees one version however often they search. When it has run, Apply offers three choices: pause the original ads and create new ads with the variation, remove the originals, or keep both. Applying ends the variation, and the new ads go through review like any others.
Google’s rule of thumb: an ad variation for one change across many campaigns, a custom experiment for several changes at once on a smaller scale.
How to read the result, and what statistically significant means
Open the experiment from Campaigns › Experiments. Its scorecard shows each metric three ways:
- The experiment’s own figure. 4K under Clicks is 4,000 clicks since it began.
- The difference from the original. +10% is 10% more than the original campaign drew.
- The range the difference may lie in. [+8%, +12%] is a confidence interval. Google shows an 80% interval unless you pick another.
A blue asterisk beside a figure marks it statistically significant. In Google’s words that means the data is “likely not due to chance”, and the experiment is more likely to keep performing the same way as a campaign. On an ad variation, Google puts a number on the asterisk: at least 95% likely that the change, and not chance, made the difference.
Read the range before the first figure. A range that runs from a small loss to a large gain hasn’t settled anything, whatever the figure beside it says.
The All experiments table sums each one up in a Results column: Control campaign where the original did better, Treatment campaign where the change did, and No clear winner or In progress where the data can’t say yet. Where a result isn’t significant, Google gives four reasons: the experiment hasn’t run long enough, the campaign doesn’t get enough traffic, the experiment’s share of traffic was too small, or the change made no measurable difference.
Don’t expect the two sides to spend the same. Google says a 50/50 split decides which auctions each side can enter, and doesn’t guarantee equal impressions or spend.
To finish, select Apply on the experiment’s page. Update your original campaign writes the change into the original. Convert to a new campaign pauses the original and runs the trial in its place, on the same budget and dates. Both keep their performance data. Some kinds of experiment apply themselves when the result is favorable. Where that exists it’s on by default, and you can turn it off from the experiment’s Report page.
Does your campaign convert enough for a test to settle anything?
Answer this before you set one up. A split test compares two counts of conversions, and small counts differ by luck. The fewer conversions the two sides share, the wider the gap between them has to be before it means anything.
Google’s page on custom experiments gives one figure: for reliable results, a base campaign should meet minimum requirements “such as more than 100 daily conversions”. Most campaigns are far smaller, so work out what yours can tell.
- Go to Campaigns › Campaigns, set the date range to the last 30 days, and read the campaign’s Conversions.
- Scale them to the length of the test. Four weeks is 28 days, so multiply by 28 and divide by 30. That is what both sides of a 50/50 test would count between them.
- In a sheet, with that count in A2, enter =EXP(2.8*SQRT(4/A2))-1. The answer is the smallest difference between the two sides the test can tell from luck: 0.2 is a fifth more conversions.
The 2.8 is the usual standard for a test. A difference that large shows up by luck 1 time in 20 when the change did nothing, and is found 4 times in 5 when the change is real.
| Conversions in four weeks | Difference it can tell | |
|---|---|---|
| 24 | 215% more | Brand — Exact at Cedar & Pine |
| 100 | 75% more | |
| 250 | 43% more | |
| 500 | 28% more | |
| 945 | 20% more | Where a fifth more shows |
| 2,800 | 11% more | 100 a day, Google’s example |
Where your campaign falls short, three things help. A longer test counts more conversions, and an experiment’s end date can be extended. A bolder change is easier to see than a fine one. And a comparison inside one ad group, next, needs no split at all.
Ad copy testing: compare two ads in one ad group
Two ads in one ad group are already a test. They share the same keywords, bids and budget, and Google asks for at least two responsive search ads in every ad group. No experiment is needed, only a fair reading, and Google warns that click-through rate or conversion rate alone can mislead. For the lines inside one ad, the ad copy guide reads each headline and description by the same sum.
- Go to Campaigns › Ads and select the Ad group column’s name to sort by it, so each ad group’s ads sit together.
- Set the date range to whole weeks in which both ads ran: the last 8 weeks, or from the newer ad’s first full week. Weeks one ad ran alone are searches the other never met.
- Read each ad’s Impr., Clicks and Conversions. To see the weeks one by one, select the segment icon, then Time, then Week. Day is offered only for a date range of 16 days or less.
- Divide each ad’s conversions by its impressions. An ad’s work starts when it’s shown, so this counts the clicks it didn’t draw as well as the clicks that didn’t buy.
- Multiply the weaker ad’s impressions by the other ad’s rate. That is what it would have brought, had it converted as often.
| Ad | Impr. | Clicks | CTR | Conversions | Conv. per impr. |
|---|---|---|---|---|---|
| Velvet Sofas | Free Delivery | 8,240 | 302 | 3.7% | 12 | 0.15% |
| Velvet Sofa Ideas | See 18 Velvet Colors | 6,180 | 228 | 3.7% | 0 | 0.00% |
The last step is how likely the gap is by luck. If both ads converted as often per impression, each of the 12 conversions had a 43% chance of following the second ad, its share of the impressions. The chance that none did is 0.57 raised to the power of 12.
| Sum | Working | Result |
|---|---|---|
| The second ad’s share of impressions | 6,180 ÷ (6,180 + 8,240) | 43% |
| Its conversions at the first ad’s rate | 6,180 × 12 ÷ 8,240 | 9 |
| Chance all 12 follow the first ad | 0.57 ^ 12 | 0.12% |
| Across the 7 ads compared | 0.12% × 7 | under 0.9% |
Set Corner sofas beside it. Its second ad brought 1 conversion from 1,960 impressions where the first brought 6 from 4,820: also under half the rate. But =BINOMDIST(1, 7, 0.29, TRUE) is 35%. A gap that size on 7 conversions happens by luck about one time in three, so pausing that ad would be a guess.
One correction is left. Compare enough ads and some gap will look convincing by luck alone, so multiply each chance by the number of ads you tested. Cedar & Pine runs two ads in four ad groups, and 7 of those 8 can be tested this way. The eighth is the Velvet sofas ad that converts, whose neighbor has no conversions to set a rate. The result is still far under 5%.
Pause the weaker ad rather than remove it: in Campaigns › Ads, select the green dot beside the ad, then Pause. A paused ad can be enabled again without being rewritten or reviewed. Then write a new ad on a different angle, so the ad group still has two to compare. Two ads that share most of their headlines test almost nothing, which the repeated headlines guide covers.
Ad rotation: what Optimize and Do not optimize each do
Ad rotation sets how often each ad in an ad group is served beside the others. No more than one ad from your account shows at a time, so an ad group’s ads take turns. The setting has two options, in Search, Shopping and Display campaigns:
- Optimize. In each auction Google favors the ads it expects to perform better, from the keyword, the search term, the device, the location and more. It optimizes for clicks, and for conversions too under a Smart Bidding strategy. As data builds, the ads likely to do better are served more often.
- Do not optimize. Ads enter the auction more evenly, for an indefinite time. Google says it “isn’t recommended for most advertisers”, and that impressions may still not come out even, since a lower-quality ad may show on a later page of results.
Under Smart Bidding the setting changes nothing. Google says that if you use Smart Bidding, Google Ads automatically uses Optimize. So in a campaign on Target CPA, Target ROAS, Maximize conversions or Maximize conversion value, your ads are never rotated evenly, whatever the setting reads.
- Go to Campaigns › Ad groups.
- Hover over the ad group, or its campaign, and select the settings icon, then Additional settings.
- Select Ad rotation, choose Optimize or Do not optimize, and select Save.
For a test, this decides what a gap between two ads means. Under Optimize, Google chose the searches each ad showed on, so part of any gap is that choice. A wide gap over several weeks is still the ad’s. Under Do not optimize, on a bid strategy that follows the setting, the ads meet comparable searches, which suits a test you read by hand.
What it looks like when Goldbeater finds it
Every day Goldbeater reads each enabled ad’s last 8 whole weeks, week by week, in every Search campaign. It holds each ad to the other ads in its ad group, over the weeks they served together. It raises an ad that converted under half as often per impression as the others, and no better per click, on at least $20 of spend, where so few conversions is unlikely by chance.
The chance it allows is 5% across every ad it tested in your account, so one ad’s gap has to be far less likely than that on its own. The finding is medium where the ad took 25% or more of its ad group’s impressions, and low under that.
An ad in Velvet sofas brought 0 conversions from 6,180 impressions, where its ad group's other ads' rate would have brought 9
Sofas — Phrase › Velvet sofas
The ad "Velvet Sofa Ideas | See 18 Velvet Colors" brought 0 conversions from 6,180 impressions over the 8 weeks it served beside another ad in Velvet sofas, where the other ad brought 12 conversions from 8,240 impressions: under a 0.9% chance of so few by luck, allowing for the 7 ads tested across the account.
- Expected at the other ads' rate
- 9
The 8 weeks they served together
| Figure | This ad (loser) | The other ad (winner) |
|---|---|---|
| Impressions | 6,180 | 8,240 |
| Clicks | 228 | 302 |
| Conversions | 0 | 12 |
| Conversions per click | 0% | 4% |
| Conversions per impression, beyond chance | 0.00% | 0.15% |
This ad
The other ad
| Velvet Sofa Ideas | See 18 Velvet Colors | 6,180 | 228 | 0 | 0.00% |
| Velvet Sofas | Free Delivery | 8,240 | 302 | 12 | 0.15% |
An ad that brought under half the conversions per impression of the other ads in its ad group over the weeks they ran together, and no more per click, by more than chance explains across every ad in the account. The impressions it takes go to the weaker message. Goldbeater drafts pausing it, and never an ad group's last ad.
Pause it, the change below, and write a new ad in Velvet sofas that takes a different angle from the one converting, so the ad group keeps two to compare. Google shows each ad on the searches it expects to convert, so the gap is partly the searches it chose; one this size is still the ad's.
- Pause the adThe ad "Velvet Sofa Ideas | See 18 Velvet Colors" in Velvet sofas
It reaches your account when you approve it in the change set and press Apply. An applied change can be undone.
This is “An ad converting far below the others in its ad group” on a sample account. Goldbeater runs it on yours every 24 hours, and your AI analyst answers what you ask about any finding.
Get a free first lookThat is the pair worked above: the finding’s 9 expected conversions and its chance are the table’s sums. The change it drafts is pausing the ad, and it never drafts a pause that would leave an ad group with no ad running. Sofas — Phrase bids to a target ROAS, so the advice says what the rotation section says: Google chose the searches for each ad, and a gap this size is still the ad’s. Goldbeater doesn’t raise Corner sofas’ second ad, for the reason the sum gave.
Goldbeater also reads your account’s experiments. Where no experiment or ad variation has run in the last 365 days, and an enabled Search campaign counted enough conversions in the last 30 days for a four-week 50/50 test to tell a fifth more, it raises one finding on the campaign counting the most. Enough is about 940 over the four weeks, or about 1,010 in the 30 days. The finding is low and is advice: run the next change to the campaign’s bidding, match types or landing pages as an experiment. It drafts no change. Cedar & Pine’s largest Search campaign counts 25.5, so nothing is raised there.
And it reads ad rotation where the setting does something: an ad group with two or more enabled ads, set to Do not optimize by its own setting or its campaign’s, in a campaign on a bid strategy outside Smart Bidding. That finding is low. It offers Optimize as a change you can choose and doesn’t recommend it, since rotating evenly is how a test read by hand is run. All three are lines of the audit checklist.
Questions
- What is an experiment in Google Ads?
- An experiment tests a change on part of a campaign’s traffic or budget while the original keeps running on the rest, so the two can be compared over the same weeks. If the change does better, you apply it to the original campaign or run the new version in its place. Experiments are set up under Campaigns, then Experiments.
- How do I A/B test ads in Google Ads?
- Three ways. Run two responsive search ads in one ad group and compare their conversions per impression over the weeks both ran. Use an ad variation to test one change of text across many ads, split by cookie. Or use a custom experiment to test a change to the campaign itself, such as its bidding or its landing pages.
- How long should a Google Ads experiment run?
- Google recommends at least 4 to 6 weeks where a result can’t yet be determined, and says to allow 7 to 14 days for the trial side to stabilize. The first end date has to be 2 to 12 weeks after the start, and you can extend it while the experiment runs.
- How many conversions does a Google Ads experiment need?
- Google’s page on custom experiments gives more than 100 daily conversions as an example of the minimum for reliable results. By the arithmetic, a four-week 50/50 test needs about 940 conversions between its two sides to tell a fifth more from luck, 250 to tell 43% more, and 100 to tell 75% more.
- Should I choose a cookie-based or a search-based split?
- Cookie-based assigns each person to one side, so nobody sees both, and Google recommends it. Search-based assigns each search, so one person can see both sides, and Google says it may reach a statistically significant result faster. A Display campaign always splits by cookie.
- What does statistically significant mean in a Google Ads experiment?
- Google marks a statistically significant figure with a blue asterisk, and says it means the data is likely not due to chance. On an ad variation, it says the asterisk means at least 95% likely that the change made the difference. Beside each figure is a confidence interval, 80% unless you pick another, giving the range the difference may lie in.
- Does ad rotation matter if I use Smart Bidding?
- No. Google says that with Smart Bidding, Google Ads automatically uses the Optimize ad rotation setting, so it favors the ads it expects to perform better whatever the setting reads. The setting changes how ads are served only under other bid strategies.