Optimization

Google Ads Experiments: How to Prove a Change Actually Worked

A before-and-after comparison bills you for seasonality and competitor moves. A campaign experiment races two versions at the same time, under the same conditions.

Google Ads experiments are the platform's built-in testing tool: they split a campaign's traffic in two and race the original version against a modified one at the same time, under the same conditions. That is the honest way to prove a change actually worked. If you compare last month with this month, what you measured is not the effect of your change; it is seasonality plus competitor moves plus your change, added together. In practice a good experiment comes down to three decisions: one variable, a 50/50 split, and the patience to wait for 95% confidence.

Measured impact: before/after vs. experimentBefore/afterExperiment316Bid change182Ad copy2411Landing page151Audience expandIllustrative scenario: change in conversions (%)
The same four changes measured two ways: the before-and-after comparison inflates the difference, the experiment shows what is really there. Illustrative scenario, not real customer data.

What are Google Ads experiments?

A Google Ads experiment is a testing framework that clones an existing campaign, sends a share of the traffic and budget to that clone, and then reports both versions side by side. You find it under the Experiments section. The original campaign is your control arm and the clone is your treatment arm. Because both run on the same days, in the same auctions, in the same season, the gap between them is largely attributable to the thing you changed.

  • Custom experiments: available for Search, Display and video campaigns, to test a bid strategy, ad copy, landing page or targeting change.
  • Performance Max experiments: test Performance Max against another campaign type, or measure the uplift it adds on top of your existing setup.
  • Asset testing: in 2026 Google extended Performance Max experimentation to compare creative assets inside a single campaign; because traffic is split within the campaign, the learning period is shorter.
  • Video and Demand Gen holdouts: reserve a control group that never sees the ad, so you can measure brand impact.

You can schedule five experiments on a single campaign, but only one runs at a time. The traffic split percentage is locked once the experiment is created; to try a different ratio you have to build a new experiment.

Why is a before-and-after comparison misleading?

Because there is no control group in it. When you compare two periods, you also compare everything else that changed between them. Even when the difference looks positive, you cannot say how much of it belongs to you. An experiment keeps a control group running in the same window, so the noise hits both arms equally and what remains is your intervention.

  • Seasonality: day of week, end of month, holidays and category season never make two periods equal.
  • Competitor moves: a rival raising budget or leaving the auction changes your costs while you do nothing; the effect shows up first in click costs.
  • Learning period: change a bid strategy and the algorithm recalibrates; the first days do not represent steady-state performance.
  • Measurement drift: conversion windows, tracking changes and data lag measure the two periods differently.
  • Regression to the mean: any action taken after a bad week looks good, because the next week was going to normalise anyway.

How do you set up a campaign experiment?

Setting one up is easier than reading one, but the setup is exactly what determines whether the reading is trustworthy. Before you start, write down the hypothesis, the primary metric and the end date. If you do not fix those up front, by the time the experiment ends you will have enough metrics to find whatever answer you wanted.

A campaign experiment in four steps1Pick onevariableChange one thingat a time2Split traffic50/50Cookie-basedsplit is best3Wait 4-6weeksFirst 7 days areexcluded4Decide at 95%Apply, revert orextend
The four steps of a campaign experiment: pick one variable, split traffic 50/50, wait at least 4-6 weeks, and decide at the 95% confidence level.
  1. Write a hypothesis: 'If I shorten the form on the landing page, conversion rate will increase.' Without a hypothesis you have curiosity, not a test.
  2. Pick the campaign and create a draft. In the draft, change only the single thing in your hypothesis.
  3. Set the traffic split to 50/50. It gives the cleanest comparison, and it cannot be changed after the experiment starts.
  4. Choose the split method. Cookie-based assignment keeps each user on one side and is the recommended option.
  5. Enter the start and end dates up front. Fixing the end date kills the reflex to look at the numbers and then move the deadline.
  6. Declare the primary metric: conversions, conversion value, CPA or ROAS. Picking whichever one looks good at the end is not an honest reading.

Start by knowing what is worth testing

Ads Sensor reads your Google, Meta, TikTok, Criteo and GA4 data in one panel and produces prioritised, reasoned actions.

Join the beta →

How long should an experiment run and how many conversions do you need?

Platform documentation recommends at least 4-6 weeks, and longer if you have a long conversion delay. But time on its own is not the criterion. The real constraint is the number of conversions accumulating in each arm, and the maths is unforgiving: the smaller the difference you want to detect, the sample size grows with the square. Halve the effect you are chasing and you need four times the data.

Conversions needed per arm6.2005% lift1.55010% lift39020% lift17030% liftIllustrative: 3% baseline CVR, 95% confidence, 80% power
The smaller the improvement you want to detect, the steeper the sample requirement. Illustrative calculation assuming a 3% baseline conversion rate, 95% confidence and 80% power.
  • At least two conversion cycles: if it takes 10 days from click to sale, a two-week experiment tells you half the story.
  • Use whole weeks: weekday and weekend behaviour differ, and a half week skews the result.
  • The first 7 days are excluded: the platform treats them as ramp-up, so week one of your experiment gives you no decision data.
  • Do not experiment on low-traffic campaigns: on a campaign with 20 conversions a month, the time needed to reach significance is longer than the time the answer stays relevant.

How do you read the result: what does significance mean?

Statistical significance means the probability that the observed difference is down to chance falls below an agreed threshold. In Google Ads experiments that threshold is a 95% confidence level, and the test is two-tailed, so improvement and deterioration are judged with the same strictness. If the reported confidence interval still contains zero, you have no result. 'It looks slightly better' is not a finding.

  • 95% confidence: the probability that the difference is random noise sits below 5%.
  • Two-tailed testing: gains and losses are held to the same bar; looking one way hides risk.
  • Resampling: the platform buckets the data into twenty groups per arm and uses jackknife resampling to estimate sample variance, which stays more conservative when data is thin.
  • No result is a result: if no difference emerged, the effect is probably small. That is not a reason to ship the change.
95%
significance threshold
4-6 weeks
recommended duration
7 days
ramp-up excluded from the report
50/50
recommended traffic split

Why one variable at a time?

Because if you change two things you will never know which one produced the result. An experiment gives you a single number: the difference between the two arms. There is no way to split that number across three changes. Worse, changes can cancel each other out; if the new copy wins while the new bid strategy loses, the total lands on zero and you throw away two good ideas at once.

  • Right: change only the bid strategy. The options are compared in our target CPA versus target ROAS guide.
  • Right: change only the landing page and leave the ad copy untouched.
  • Wrong: raise the budget and broaden match types in the same experiment.
  • Wrong: add negative keywords to the original campaign while the experiment is running. Edit the control mid-flight and the comparison is void.
  • Creative needs its own frame: to test image and video variants in sequence, see the creative testing framework; same logic, different scale.

What are the six most common mistakes?

Most experiments do not break during setup, they break during the reading. These six mistakes make the majority of results unusable in practice.

  1. Deciding too early: declaring the arm that is ahead on day five. Checking constantly and stopping at the first good moment manufactures false positives; a fixed end date prevents it.
  2. Testing in a seasonal window: Black Friday, a holiday or a big promo week changes user behaviour. The answer belongs to that week and does not transfer to the rest of the year.
  3. Overlapping experiments: running several tests at once in the same account lets them interfere with each other and makes the results less reliable. Run them sequentially.
  4. Choosing the metric afterwards: when conversions show nothing, switching to click-through rate. Look at enough metrics and one of them will come back significant.
  5. Ending the experiment manually: experiments that finish on their scheduled date can have a favourable result applied automatically; manually ended ones are excluded from that.
  6. Generalising the result: a setting that wins in one campaign may lose across the account. A bid strategy that works on brand terms behaves differently on generic ones, a split we also unpack in the quality score guide.

Where does Ads Sensor fit in this loop?

Ads Sensor does not set up the experiment for you; that stays in the Google Ads interface. Its contribution sits before and after. Before, it reads your Google, Meta, TikTok, Criteo and GA4 data in one place and works out which campaign is worth testing and what evidence your hypothesis rests on. After, it runs automatic before-and-after result tracking on the recommendations you approve and apply from the panel. To be straight about it: that tracking is a record of what changed, not proof of causation. For causation you still need an experiment. You can join the beta and connect your accounts.

Frequently asked questions

Do Google Ads experiments need extra budget?
No. Custom experiments share the original campaign's traffic and budget. With a 50% split, roughly half the spend goes to the experiment arm. Your total spend stays the same, it is just divided in two.
What is the minimum duration for an experiment?
In practice 4-6 weeks is a good floor. If your conversion delay is long, wait for at least two conversion cycles. Because the first 7 days are treated as ramp-up and excluded, a two-week experiment effectively gives you one week of usable data.
What should I do if the result is not significant?
You have three options: extend the experiment, repeat it on a higher-volume campaign, or accept that the difference is genuinely small and move to another hypothesis. Reading an inconclusive result as 'let's ship it anyway' puts you back where you started.
Can I run several experiments at the same time?
Technically yes, but it is not recommended. Experiments interfere with each other's traffic and learning, which makes results less reliable. You can schedule five experiments on one campaign but only one runs at a time, and account-wide it is safer to go sequentially.
What is the difference between cookie-based and search-based splits?
With a cookie-based split a user always stays in the same arm. With a search-based split the assignment is redrawn on every search, so one person can see both versions. Search-based can reach significance faster; cookie-based gives a cleaner comparison and is the recommended option.
The experiment won, what now?
Experiments that end on their scheduled date can have a favourable result applied automatically. You can also apply the change to the original campaign manually or convert the experiment campaign into a permanent one. Watch performance for at least a month afterwards: a winning setting can behave differently at scale.

Find what is worth testing

Ads Sensor unifies five data sources in one panel and lists risks and opportunities with the reasoning behind them. You decide which hypothesis earns an experiment.

Join the beta →