Google Ads experiments are the platform's built-in testing tool: they split a campaign's traffic in two and race the original version against a modified one at the same time, under the same conditions. That is the honest way to prove a change actually worked. If you compare last month with this month, what you measured is not the effect of your change; it is seasonality plus competitor moves plus your change, added together. In practice a good experiment comes down to three decisions: one variable, a 50/50 split, and the patience to wait for 95% confidence.
What are Google Ads experiments?
A Google Ads experiment is a testing framework that clones an existing campaign, sends a share of the traffic and budget to that clone, and then reports both versions side by side. You find it under the Experiments section. The original campaign is your control arm and the clone is your treatment arm. Because both run on the same days, in the same auctions, in the same season, the gap between them is largely attributable to the thing you changed.
- Custom experiments: available for Search, Display and video campaigns, to test a bid strategy, ad copy, landing page or targeting change.
- Performance Max experiments: test Performance Max against another campaign type, or measure the uplift it adds on top of your existing setup.
- Asset testing: in 2026 Google extended Performance Max experimentation to compare creative assets inside a single campaign; because traffic is split within the campaign, the learning period is shorter.
- Video and Demand Gen holdouts: reserve a control group that never sees the ad, so you can measure brand impact.
You can schedule five experiments on a single campaign, but only one runs at a time. The traffic split percentage is locked once the experiment is created; to try a different ratio you have to build a new experiment.
Why is a before-and-after comparison misleading?
Because there is no control group in it. When you compare two periods, you also compare everything else that changed between them. Even when the difference looks positive, you cannot say how much of it belongs to you. An experiment keeps a control group running in the same window, so the noise hits both arms equally and what remains is your intervention.
- Seasonality: day of week, end of month, holidays and category season never make two periods equal.
- Competitor moves: a rival raising budget or leaving the auction changes your costs while you do nothing; the effect shows up first in click costs.
- Learning period: change a bid strategy and the algorithm recalibrates; the first days do not represent steady-state performance.
- Measurement drift: conversion windows, tracking changes and data lag measure the two periods differently.
- Regression to the mean: any action taken after a bad week looks good, because the next week was going to normalise anyway.
How do you set up a campaign experiment?
Setting one up is easier than reading one, but the setup is exactly what determines whether the reading is trustworthy. Before you start, write down the hypothesis, the primary metric and the end date. If you do not fix those up front, by the time the experiment ends you will have enough metrics to find whatever answer you wanted.
- Write a hypothesis: 'If I shorten the form on the landing page, conversion rate will increase.' Without a hypothesis you have curiosity, not a test.
- Pick the campaign and create a draft. In the draft, change only the single thing in your hypothesis.
- Set the traffic split to 50/50. It gives the cleanest comparison, and it cannot be changed after the experiment starts.
- Choose the split method. Cookie-based assignment keeps each user on one side and is the recommended option.
- Enter the start and end dates up front. Fixing the end date kills the reflex to look at the numbers and then move the deadline.
- Declare the primary metric: conversions, conversion value, CPA or ROAS. Picking whichever one looks good at the end is not an honest reading.
Start by knowing what is worth testing
Ads Sensor reads your Google, Meta, TikTok, Criteo and GA4 data in one panel and produces prioritised, reasoned actions.
How long should an experiment run and how many conversions do you need?
Platform documentation recommends at least 4-6 weeks, and longer if you have a long conversion delay. But time on its own is not the criterion. The real constraint is the number of conversions accumulating in each arm, and the maths is unforgiving: the smaller the difference you want to detect, the sample size grows with the square. Halve the effect you are chasing and you need four times the data.
- At least two conversion cycles: if it takes 10 days from click to sale, a two-week experiment tells you half the story.
- Use whole weeks: weekday and weekend behaviour differ, and a half week skews the result.
- The first 7 days are excluded: the platform treats them as ramp-up, so week one of your experiment gives you no decision data.
- Do not experiment on low-traffic campaigns: on a campaign with 20 conversions a month, the time needed to reach significance is longer than the time the answer stays relevant.
How do you read the result: what does significance mean?
Statistical significance means the probability that the observed difference is down to chance falls below an agreed threshold. In Google Ads experiments that threshold is a 95% confidence level, and the test is two-tailed, so improvement and deterioration are judged with the same strictness. If the reported confidence interval still contains zero, you have no result. 'It looks slightly better' is not a finding.
- 95% confidence: the probability that the difference is random noise sits below 5%.
- Two-tailed testing: gains and losses are held to the same bar; looking one way hides risk.
- Resampling: the platform buckets the data into twenty groups per arm and uses jackknife resampling to estimate sample variance, which stays more conservative when data is thin.
- No result is a result: if no difference emerged, the effect is probably small. That is not a reason to ship the change.
Why one variable at a time?
Because if you change two things you will never know which one produced the result. An experiment gives you a single number: the difference between the two arms. There is no way to split that number across three changes. Worse, changes can cancel each other out; if the new copy wins while the new bid strategy loses, the total lands on zero and you throw away two good ideas at once.
- Right: change only the bid strategy. The options are compared in our target CPA versus target ROAS guide.
- Right: change only the landing page and leave the ad copy untouched.
- Wrong: raise the budget and broaden match types in the same experiment.
- Wrong: add negative keywords to the original campaign while the experiment is running. Edit the control mid-flight and the comparison is void.
- Creative needs its own frame: to test image and video variants in sequence, see the creative testing framework; same logic, different scale.
What are the six most common mistakes?
Most experiments do not break during setup, they break during the reading. These six mistakes make the majority of results unusable in practice.
- Deciding too early: declaring the arm that is ahead on day five. Checking constantly and stopping at the first good moment manufactures false positives; a fixed end date prevents it.
- Testing in a seasonal window: Black Friday, a holiday or a big promo week changes user behaviour. The answer belongs to that week and does not transfer to the rest of the year.
- Overlapping experiments: running several tests at once in the same account lets them interfere with each other and makes the results less reliable. Run them sequentially.
- Choosing the metric afterwards: when conversions show nothing, switching to click-through rate. Look at enough metrics and one of them will come back significant.
- Ending the experiment manually: experiments that finish on their scheduled date can have a favourable result applied automatically; manually ended ones are excluded from that.
- Generalising the result: a setting that wins in one campaign may lose across the account. A bid strategy that works on brand terms behaves differently on generic ones, a split we also unpack in the quality score guide.
Where does Ads Sensor fit in this loop?
Ads Sensor does not set up the experiment for you; that stays in the Google Ads interface. Its contribution sits before and after. Before, it reads your Google, Meta, TikTok, Criteo and GA4 data in one place and works out which campaign is worth testing and what evidence your hypothesis rests on. After, it runs automatic before-and-after result tracking on the recommendations you approve and apply from the panel. To be straight about it: that tracking is a record of what changed, not proof of causation. For causation you still need an experiment. You can join the beta and connect your accounts.