Guide

How do Google Ads campaign experiments and A/B testing work?

The kinds of Google Ads experiments, how to set one up and read its result, how many conversions a test needs, and ad copy testing without an experiment.

In short
  • A Google Ads experiment runs a change on part of a campaign's traffic or budget beside the original, over the same weeks, and compares the two. You set one up under Campaigns, Experiments.
  • A custom experiment tests bidding, match types, landing pages or audiences on a trial copy of a campaign, and an ad variation tests one change of ad text across campaigns. Google recommends a 50% split by cookie, and at least 4 to 6 weeks where a result can't yet be determined.
  • Google marks a statistically significant figure with a blue asterisk and shows the range the difference may lie in, an 80% confidence interval unless you pick another.
  • A test settles only what its conversions allow. A four-week 50/50 test needs about 940 conversions to tell a fifth more from luck, and Google's own example of enough is more than 100 a day.
  • Goldbeater compares the ads in every ad group every day, and raises one that converts under half as often per impression as the others by more than chance explains, with its pause drafted.
Goldbeater checks this every 24 hours
  • An ad converting far below the others in its ad group
  • A campaign big enough to test changes that never runs an experiment
  • Ads rotating evenly, however each performs

Each week Goldbeater's AI agent follows up on the findings. A free first look shows how many problems these and 160 other checks find on one of your accounts, and what they cost. No card.

The manual way comes first, in Google Ads itself. Then the same problem as Goldbeater finds it, every day, without anyone asking.

Google Ads A/B testing splits one campaign and compares the halves.

An experiment runs a change beside the campaign it would change. Google splits the campaign’s traffic or budget between the original and a version with your change in it, both run over the same weeks, and you compare them at the end. If the change did better, you apply it to the original campaign or run the new version in its place.

That is an A/B test, and it answers what a before-and-after can’t. A season, a sale or a new competitor moves both halves at once, so what’s left between them is the change.

Experiments are set up under Campaigns › Experiments, a page Google used to call Drafts and experiments. All experiments lists every one in the account. Google’s page names seven kinds today:

KindWhat it testsCampaigns
Custom experimentSmart Bidding, keyword match types, landing pages, audiences and ad groups, in a trial copy of the campaignSearch, Display, Video
Ad variationOne change of ad text or URL across many ads, such as one phrase swapped for anotherSearch
AI Max experimentAI Max’s search term matching and asset optimization, on half of one campaign, with no copy madeSearch
Performance Max experimentWhat Performance Max adds beside your other campaigns, a move to it from Standard Shopping, or a setting or asset inside itPerformance Max
Video experimentWhich of your video ads is more effective on YouTubeVideo
Demand Gen experimentWhich images or videos draw the most interestDemand Gen
App asset experimentThe text and images inside an App campaignApp

Two of them cover most of what a Search advertiser tests. A custom experiment tests a change to the campaign itself: how it bids, what it matches, where it lands. An ad variation tests a change to ad text, and Google recommends it over a custom experiment for Search ads. Google’s pages on custom experiments differ on Demand Gen and Hotel Ads campaigns, and agree that an App or a Shopping campaign can’t take one. Performance Max experiments have a set of their own: uplift, upgrade and optimization.

How to set up a Google Ads campaign experiment

First check the campaign can take a custom experiment. Google lists what it needs:

  • Its own budget and bid strategy. The campaign has to be enabled, on an individual daily budget and a standard bid strategy. One on a shared budget or a portfolio bid strategy may not appear in the list.
  • No expanded text ads. Google can’t create an experiment for a campaign that holds expanded text ads or older text ads, in paused and removed ad groups too. Remove them and allow 24 to 48 hours. The expanded text ads guide shows how to find them.
  • No other experiment on those dates. A campaign can have five scheduled, and only one running at a time.
  1. Go to Campaigns › Experiments, select the plus button, then Custom.
  2. Select Custom experiment and the campaign type, then Continue.
  3. Name the experiment, choose the campaign to test, and give its copy a suffix. Google creates the trial campaign from the original.
  4. Make the change you’re testing in the trial campaign, and choose up to two goals to measure it by. Google’s example is Clicks and Increase.
  5. Under Experiment split, set the share of traffic and budget the trial takes. Google recommends 50%, for the best comparison between the two.
  6. Under Advanced, choose how people are assigned: Cookie-based or Search-based.
  7. Under Schedule the experiment, set the start and the duration, then select Save.

There’s a shorter way in: edit an enabled campaign’s settings and select Save as experiment, which builds the trial from the change you just made.

  • Cookie-based assigns each person to one side, so nobody sees both. Google recommends it. If the campaign uses audience lists, Google says each should hold at least 10,000 people, or the results may be less accurate.
  • Search-based assigns each search, so one person can meet both sides on different searches. Google says it may reach a statistically significant result faster.
  • A Display campaign always splits by cookie.

Set it to run at least four weeks. Google recommends 4 to 6 weeks where a result can’t yet be determined, and says to allow 7 to 14 days for the trial side to stabilize. A trial that stops serving may be in its bid strategy’s learning phase, typically 7 days, which the learning phase guide covers. The first end date has to fall 2 to 12 weeks after the start, and you can extend it while the experiment runs. Once it ends, it can’t be resumed.

Then leave both sides alone. The split can’t be changed once the experiment is set up, changes to the original don’t reach the trial, and Google says changes to either side make the result harder to read. It also advises against running several experiments at once, since they can interfere with each other.

Google Ads ad variations test one change of ad text across campaigns

An ad variation changes the same words in many ads at once and shows the changed ads to a share of people. Google’s example changes a call to action from “Buy now” to “Buy today”. It works on responsive search ads.

  1. Go to Campaigns › Experiments, open the Ad variations tab and select the plus button.
  2. Under Select ads, choose All campaigns or one campaign, and under Target ad type, Responsive search ads. Filter ads narrows it by headline, description, display path or final URL.
  3. Under Create variation, choose Find and replace, Update text or Update URLs. Variations are case sensitive.
  4. Name it, set its start and end dates, and under Experiment split enter the share of the campaign’s budget it takes. Then select Create variation.

People are split by cookie, so a person sees one version however often they search. When it has run, Apply offers three choices: pause the original ads and create new ads with the variation, remove the originals, or keep both. Applying ends the variation, and the new ads go through review like any others.

Google’s rule of thumb: an ad variation for one change across many campaigns, a custom experiment for several changes at once on a smaller scale.

How to read the result, and what statistically significant means

Open the experiment from Campaigns › Experiments. Its scorecard shows each metric three ways:

  • The experiment’s own figure. 4K under Clicks is 4,000 clicks since it began.
  • The difference from the original. +10% is 10% more than the original campaign drew.
  • The range the difference may lie in. [+8%, +12%] is a confidence interval. Google shows an 80% interval unless you pick another.

A blue asterisk beside a figure marks it statistically significant. In Google’s words that means the data is “likely not due to chance”, and the experiment is more likely to keep performing the same way as a campaign. On an ad variation, Google puts a number on the asterisk: at least 95% likely that the change, and not chance, made the difference.

Read the range before the first figure. A range that runs from a small loss to a large gain hasn’t settled anything, whatever the figure beside it says.

The All experiments table sums each one up in a Results column: Control campaign where the original did better, Treatment campaign where the change did, and No clear winner or In progress where the data can’t say yet. Where a result isn’t significant, Google gives four reasons: the experiment hasn’t run long enough, the campaign doesn’t get enough traffic, the experiment’s share of traffic was too small, or the change made no measurable difference.

Don’t expect the two sides to spend the same. Google says a 50/50 split decides which auctions each side can enter, and doesn’t guarantee equal impressions or spend.

To finish, select Apply on the experiment’s page. Update your original campaign writes the change into the original. Convert to a new campaign pauses the original and runs the trial in its place, on the same budget and dates. Both keep their performance data. Some kinds of experiment apply themselves when the result is favorable. Where that exists it’s on by default, and you can turn it off from the experiment’s Report page.

Does your campaign convert enough for a test to settle anything?

Answer this before you set one up. A split test compares two counts of conversions, and small counts differ by luck. The fewer conversions the two sides share, the wider the gap between them has to be before it means anything.

Google’s page on custom experiments gives one figure: for reliable results, a base campaign should meet minimum requirements “such as more than 100 daily conversions”. Most campaigns are far smaller, so work out what yours can tell.

  1. Go to Campaigns › Campaigns, set the date range to the last 30 days, and read the campaign’s Conversions.
  2. Scale them to the length of the test. Four weeks is 28 days, so multiply by 28 and divide by 30. That is what both sides of a 50/50 test would count between them.
  3. In a sheet, with that count in A2, enter =EXP(2.8*SQRT(4/A2))-1. The answer is the smallest difference between the two sides the test can tell from luck: 0.2 is a fifth more conversions.

The 2.8 is the usual standard for a test. A difference that large shows up by luck 1 time in 20 when the change did nothing, and is found 4 times in 5 when the change is real.

A four-week test
Example
Conversions in four weeksDifference it can tell
24215% moreBrand — Exact at Cedar & Pine
10075% more
25043% more
50028% more
94520% moreWhere a fifth more shows
2,80011% more100 a day, Google’s example
Brand — Exact, the largest Search campaign of the sample account Cedar & Pine, counted 25.5 conversions in 30 days, about 24 in four weeks. A test there could only tell a change that more than tripled its conversions. At about 940 conversions, a fifth more shows.

Where your campaign falls short, three things help. A longer test counts more conversions, and an experiment’s end date can be extended. A bolder change is easier to see than a fine one. And a comparison inside one ad group, next, needs no split at all.

Ad copy testing: compare two ads in one ad group

Two ads in one ad group are already a test. They share the same keywords, bids and budget, and Google asks for at least two responsive search ads in every ad group. No experiment is needed, only a fair reading, and Google warns that click-through rate or conversion rate alone can mislead. For the lines inside one ad, the ad copy guide reads each headline and description by the same sum.

  1. Go to Campaigns › Ads and select the Ad group column’s name to sort by it, so each ad group’s ads sit together.
  2. Set the date range to whole weeks in which both ads ran: the last 8 weeks, or from the newer ad’s first full week. Weeks one ad ran alone are searches the other never met.
  3. Read each ad’s Impr., Clicks and Conversions. To see the weeks one by one, select the segment icon, then Time, then Week. Day is offered only for a date range of 16 days or less.
  4. Divide each ad’s conversions by its impressions. An ad’s work starts when it’s shown, so this counts the clicks it didn’t draw as well as the clicks that didn’t buy.
  5. Multiply the weaker ad’s impressions by the other ad’s rate. That is what it would have brought, had it converted as often.
Velvet sofas, 8 weeks
Sample data
AdImpr.ClicksCTRConversionsConv. per impr.
Velvet Sofas | Free Delivery8,2403023.7%120.15%
Velvet Sofa Ideas | See 18 Velvet Colors6,1802283.7%00.00%
Sofas — Phrase › Velvet sofas at Cedar & Pine. Both ads draw clicks from 3.7% of their impressions, so by click-through rate they are level. Every sale followed the first ad.

The last step is how likely the gap is by luck. If both ads converted as often per impression, each of the 12 conversions had a 43% chance of following the second ad, its share of the impressions. The chance that none did is 0.57 raised to the power of 12.

The gap, worked
Sample data
SumWorkingResult
The second ad’s share of impressions6,180 ÷ (6,180 + 8,240)43%
Its conversions at the first ad’s rate6,180 × 12 ÷ 8,2409
Chance all 12 follow the first ad0.57 ^ 120.12%
Across the 7 ads compared0.12% × 7under 0.9%
About 1 time in 800, the second ad would bring nothing by luck. Where the weaker ad has some conversions, a sheet does the same sum: =BINOMDIST(its, both, share, TRUE), with its conversions, both ads’ conversions and its share of impressions.

Set Corner sofas beside it. Its second ad brought 1 conversion from 1,960 impressions where the first brought 6 from 4,820: also under half the rate. But =BINOMDIST(1, 7, 0.29, TRUE) is 35%. A gap that size on 7 conversions happens by luck about one time in three, so pausing that ad would be a guess.

One correction is left. Compare enough ads and some gap will look convincing by luck alone, so multiply each chance by the number of ads you tested. Cedar & Pine runs two ads in four ad groups, and 7 of those 8 can be tested this way. The eighth is the Velvet sofas ad that converts, whose neighbor has no conversions to set a rate. The result is still far under 5%.

Pause the weaker ad rather than remove it: in Campaigns › Ads, select the green dot beside the ad, then Pause. A paused ad can be enabled again without being rewritten or reviewed. Then write a new ad on a different angle, so the ad group still has two to compare. Two ads that share most of their headlines test almost nothing, which the repeated headlines guide covers.

Ad rotation: what Optimize and Do not optimize each do

Ad rotation sets how often each ad in an ad group is served beside the others. No more than one ad from your account shows at a time, so an ad group’s ads take turns. The setting has two options, in Search, Shopping and Display campaigns:

  • Optimize. In each auction Google favors the ads it expects to perform better, from the keyword, the search term, the device, the location and more. It optimizes for clicks, and for conversions too under a Smart Bidding strategy. As data builds, the ads likely to do better are served more often.
  • Do not optimize. Ads enter the auction more evenly, for an indefinite time. Google says it “isn’t recommended for most advertisers”, and that impressions may still not come out even, since a lower-quality ad may show on a later page of results.

Under Smart Bidding the setting changes nothing. Google says that if you use Smart Bidding, Google Ads automatically uses Optimize. So in a campaign on Target CPA, Target ROAS, Maximize conversions or Maximize conversion value, your ads are never rotated evenly, whatever the setting reads.

  1. Go to Campaigns › Ad groups.
  2. Hover over the ad group, or its campaign, and select the settings icon, then Additional settings.
  3. Select Ad rotation, choose Optimize or Do not optimize, and select Save.

For a test, this decides what a gap between two ads means. Under Optimize, Google chose the searches each ad showed on, so part of any gap is that choice. A wide gap over several weeks is still the ad’s. Under Do not optimize, on a bid strategy that follows the setting, the ads meet comparable searches, which suits a test you read by hand.

What it looks like when Goldbeater finds it

Every day Goldbeater reads each enabled ad’s last 8 whole weeks, week by week, in every Search campaign. It holds each ad to the other ads in its ad group, over the weeks they served together. It raises an ad that converted under half as often per impression as the others, and no better per click, on at least $20 of spend, where so few conversions is unlikely by chance.

The chance it allows is 5% across every ad it tested in your account, so one ad’s gap has to be far less likely than that on its own. The finding is medium where the ad took 25% or more of its ad group’s impressions, and low under that.

Finding
Sample data
MediumMedium severityConversions

An ad in Velvet sofas brought 0 conversions from 6,180 impressions, where its ad group's other ads' rate would have brought 9

Sofas — Phrase › Velvet sofas

What Goldbeater saw

The ad "Velvet Sofa Ideas | See 18 Velvet Colors" brought 0 conversions from 6,180 impressions over the 8 weeks it served beside another ad in Velvet sofas, where the other ad brought 12 conversions from 8,240 impressions: under a 0.9% chance of so few by luck, allowing for the 7 ads tested across the account.

Expected at the other ads' rate
9

The 8 weeks they served together

FigureThis ad (loser)The other ad (winner)
Impressions6,1808,240
Clicks228302
Conversions012
Conversions per click0%4%
Conversions per impression, beyond chance0.00%0.15%
Loser

This ad

Winner

The other ad

Velvet Sofa Ideas | See 18 Velvet Colors6,18022800.00%
Velvet Sofas | Free Delivery8,240302120.15%
Why it matters

An ad that brought under half the conversions per impression of the other ads in its ad group over the weeks they ran together, and no more per click, by more than chance explains across every ad in the account. The impressions it takes go to the weaker message. Goldbeater drafts pausing it, and never an ad group's last ad.

What to do

Pause it, the change below, and write a new ad in Velvet sofas that takes a different angle from the one converting, so the ad group keeps two to compare. Google shows each ad on the searches it expects to convert, so the gap is partly the searches it chose; one this size is still the ad's.

What would change
  • Pause the adThe ad "Velvet Sofa Ideas | See 18 Velvet Colors" in Velvet sofas

It reaches your account when you approve it in the change set and press Apply. An applied change can be undone.

This is “An ad converting far below the others in its ad group” on a sample account. Goldbeater runs it on yours every 24 hours, and your AI analyst answers what you ask about any finding.

Get a free first look
Sample data. Cedar & Pine Interiors is an invented account, and its figures are made up.

That is the pair worked above: the finding’s 9 expected conversions and its chance are the table’s sums. The change it drafts is pausing the ad, and it never drafts a pause that would leave an ad group with no ad running. Sofas — Phrase bids to a target ROAS, so the advice says what the rotation section says: Google chose the searches for each ad, and a gap this size is still the ad’s. Goldbeater doesn’t raise Corner sofas’ second ad, for the reason the sum gave.

Goldbeater also reads your account’s experiments. Where no experiment or ad variation has run in the last 365 days, and an enabled Search campaign counted enough conversions in the last 30 days for a four-week 50/50 test to tell a fifth more, it raises one finding on the campaign counting the most. Enough is about 940 over the four weeks, or about 1,010 in the 30 days. The finding is low and is advice: run the next change to the campaign’s bidding, match types or landing pages as an experiment. It drafts no change. Cedar & Pine’s largest Search campaign counts 25.5, so nothing is raised there.

And it reads ad rotation where the setting does something: an ad group with two or more enabled ads, set to Do not optimize by its own setting or its campaign’s, in a campaign on a bid strategy outside Smart Bidding. That finding is low. It offers Optimize as a change you can choose and doesn’t recommend it, since rotating evenly is how a test read by hand is run. All three are lines of the audit checklist.

Questions

What is an experiment in Google Ads?
An experiment tests a change on part of a campaign’s traffic or budget while the original keeps running on the rest, so the two can be compared over the same weeks. If the change does better, you apply it to the original campaign or run the new version in its place. Experiments are set up under Campaigns, then Experiments.
How do I A/B test ads in Google Ads?
Three ways. Run two responsive search ads in one ad group and compare their conversions per impression over the weeks both ran. Use an ad variation to test one change of text across many ads, split by cookie. Or use a custom experiment to test a change to the campaign itself, such as its bidding or its landing pages.
How long should a Google Ads experiment run?
Google recommends at least 4 to 6 weeks where a result can’t yet be determined, and says to allow 7 to 14 days for the trial side to stabilize. The first end date has to be 2 to 12 weeks after the start, and you can extend it while the experiment runs.
How many conversions does a Google Ads experiment need?
Google’s page on custom experiments gives more than 100 daily conversions as an example of the minimum for reliable results. By the arithmetic, a four-week 50/50 test needs about 940 conversions between its two sides to tell a fifth more from luck, 250 to tell 43% more, and 100 to tell 75% more.
Should I choose a cookie-based or a search-based split?
Cookie-based assigns each person to one side, so nobody sees both, and Google recommends it. Search-based assigns each search, so one person can see both sides, and Google says it may reach a statistically significant result faster. A Display campaign always splits by cookie.
What does statistically significant mean in a Google Ads experiment?
Google marks a statistically significant figure with a blue asterisk, and says it means the data is likely not due to chance. On an ad variation, it says the asterisk means at least 95% likely that the change made the difference. Beside each figure is a confidence interval, 80% unless you pick another, giving the range the difference may lie in.
Does ad rotation matter if I use Smart Bidding?
No. Google says that with Smart Bidding, Google Ads automatically uses the Optimize ad rotation setting, so it favors the ads it expects to perform better whatever the setting reads. The setting changes how ads are served only under other bid strategies.

How Goldbeater checks this

These checks run every day, and whenever you ask.

Each finding shows its evidence, an estimate of what it costs a month where it costs money, and, where a setting fixes it, a change drafted for you to approve.

A check is code that runs the same way every time, so a new finding means your account changed. Each week Goldbeater's AI agent follows up on what the checks find, and your AI analyst answers any question about your account.
ChecksA finding is worth
  • An ad converting far below the others in its ad groupFixing it raises the Conversions
  • A campaign big enough to test changes that never runs an experimentFixing it raises the Conversions
  • Ads rotating evenly, however each performsYour call

7-day trial · nothing charged for 7 days

Try it for a week. See the money it saves. It pays for itself.

Start the 7-day trial

Or connect one account for a free first look, no card needed: how many problems it finds, and what they cost.