Experiments in Google Ads 2026: Testing Hypotheses in Experiment Center
“We switched the bid strategy and things improved.” Two weeks later you learn that on the same Monday a competitor cut their budget, and marketing sent a promotional email. What actually caused the lift? Nobody knows — but the decision has already been rolled out across the account. Experiments in Google Ads exist to prevent exactly this: they split the same traffic, serve two versions of the same campaign, and measure the difference under identical external conditions.
In January 2026 Google consolidated its scattered testing tools into a single Experiment Center. Here’s what’s in it, how to split traffic properly, how long to wait, which hypotheses are testable this way and which are not, and how to read a result without mistaking noise for a win.
How experiments in Google Ads work: drafts, splits, significance
The mechanism is simple, which is why it’s reliable. You create a draft — an exact copy of a campaign that you can edit freely without touching the original. From the draft you launch an experiment, and Google begins randomly dividing the campaign’s incoming traffic between two versions: control (the original) and treatment (your draft).
The critical difference from before-and-after analysis is that both versions run simultaneously. Seasonality, weather, competitor behavior, your email blast, currency moves — all of it hits both groups equally and cancels out in the comparison. That’s why an experiment answers “what did this change do,” and period-over-period comparison never does.
The split happens at the user level, not the query level: the same person consistently lands in the same group. That matters for longer consideration cycles — otherwise a user would bounce between versions and the effect would wash out.
What Experiment Center is and what changed in 2026
Testing used to be scattered across the interface: drafts and experiments in one place, ad variation tests in another, lift studies only through a rep. In January 2026 Google pulled it together into Experiment Center, a single hub housing two fundamentally different classes of measurement:
- Experiments — classic A/B tests with a traffic split inside a campaign. They answer “which setting performs better.”
- Lift studies — incrementality measurement, where a holdout group sees no ads at all and behavior is compared against the exposed group. They answer “did the advertising add anything.”
Within the hub, experiments come in four flavors:
| Experiment type | What it tests | Typical question |
|---|---|---|
| Campaign features and settings | Bid strategies, budget, targeting, ad scheduling | Is tROAS 400% better than tCPA $40? |
| Assets | Headlines, descriptions, images, video | Does the new creative set convert better? |
| Campaign types | The contribution of adding a campaign type to the mix | Did Demand Gen add sales or cannibalize search? |
| Custom experiments | Multiple variables, including across campaigns | Does the new structure beat the old one overall? |
Campaign type coverage is broad: Search, Display, Shopping, Video, App, Demand Gen, and Performance Max. Some advanced setups — particularly lift studies with heavy volume requirements — still need coordination with a Google account representative and can’t simply be toggled on in a small account.
Launching an experiment, step by step
- Write the hypothesis down before you build the draft. Format: “If we [change], then [metric] moves by [magnitude], because [mechanism].” A hypothesis without an expected magnitude isn’t a hypothesis — it’s curiosity, and you’ll rationalize any outcome after the fact.
- Create the draft and make exactly one change. Two changes mean two competing explanations for any result.
- Convert the draft into an experiment, setting dates and split share.
- Choose the split. The 50/50 default is right in most situations.
- Fix the duration up front. Two to four weeks depending on conversion volume — and never fewer than two complete weekly cycles.
- Change nothing mid-flight. Editing inside a running experiment destroys its meaning: you no longer know which version produced what.
- Wait for significance, then decide. Google flags a result as statistically significant at 95% confidence — meaning under a 5% probability the observed difference is random.
- Apply or discard. A winning variant can be applied back to the original campaign in one action, or split off into its own campaign.
Why 50/50 rather than 10/90
“Let’s test carefully, on 10% of traffic” is an understandable instinct and almost always the wrong call. Test sensitivity depends on the conversion volume in the smaller group. At 10% you’ll accumulate the required statistics roughly ten times slower — and during that stretch everything the experiment was supposed to control for will have changed. A small share makes sense in exactly one case: the change is risky and potentially expensive, and you’re consciously paying in duration for safety.
How much data you actually need
A proper sample size calculation needs an expected effect and a baseline rate, but a rough anchor is enough for planning. The smaller the effect you want to detect, the more conversions you need — and the relationship is quadratic: halving the detectable effect quadruples the data requirement.
| Expected effect | Rough conversions per group | Practical conclusion |
|---|---|---|
| +30% or more | ~100–150 | Detectable in two weeks in a mid-sized account |
| +15% | ~400–600 | Needs real volume or a month of runtime |
| +5% | ~3,000–5,000 | Out of reach for most accounts — don’t test it |
Those figures are orders of magnitude, not a calculation. But the conclusion is hard and practical: if a campaign produces 20 conversions a week, testing 5% improvements is pointless. Accounts at that scale should test large structural hypotheses — a different bid strategy, a genuinely different landing page, a different account structure — and simply implement small refinements on judgment without torturing the statistics.
What to test and what not to test
Good candidates:
- bid strategy changes or meaningful target shifts (tCPA → tROAS, a 20–30% move in target) — the trade-offs are covered in the breakdown of Google Ads bidding strategies;
- match type expansion, where intuition fails more often than anywhere else — see the guide to keyword match types;
- a new landing page against the current one;
- a different ad group structure (consolidation versus granularity);
- geo expansion or contraction and bid adjustments;
- adding or removing a campaign type from the mix.
Poor candidates:
- Small copy tweaks inside RSAs. Google rotates combinations itself, and the effect drowns at campaign level. Ad-level testing follows different logic — covered in the piece on RSA ad A/B testing.
- Incrementality questions. “Does brand search add sales” cannot be answered by an in-campaign split — that needs a geo experiment or a lift study, as described in incrementality and geo experiments.
- Anything that fixes an obvious defect. A broken conversion tag or a flood of irrelevant queries isn’t a hypothesis, it’s a bug. Fix it now.
- Changes on tiny volume. See the table above.
Reading results without fooling yourself
- Peeking. Checking daily and stopping the moment the variant pulls ahead is the fastest way to systematically ship noise. Set the end date in advance and don’t move it because “you can already see it.”
- Winning on the wrong metric. More conversions with lower conversion value is not a win. Judge on the metric tied to money, not the one that feels best.
- Learning period inside the test. After a bid strategy change the variant underperforms for the first 7–14 days simply because it’s learning. Either build learning into the runtime and discard the first week in analysis, or extend the test.
- Conversion lag. If deals close over ten days, the last ten days of the experiment are undercounted in both groups. Analyze with a delay equal to your typical lag.
- Significance without magnitude. A statistically significant 1.5% lift on high volume is real and useless. Ask not only “is it significant” but “is it worth shipping.”
An experiment doesn’t answer “did things get better.” It answers “did things get better because of our change.” That’s the only answer you can safely generalize to other campaigns.
Testing as a process, not an event
Accounts where testing pays off differ not in tooling but in discipline. The working setup looks like this:
- A hypothesis backlog. A simple table: hypothesis, expected effect, implementation cost, source (search terms report, audit, a Google recommendation). Prioritize by effect divided by effort.
- One or two active experiments per campaign. More than that and you lose the ability to attribute anything.
- A results log — including failures. Re-running the same hypothesis six months later is the most common waste of time in ad teams.
- A recurring source of hypotheses. Search terms reports, the quarterly audit, and the recommendations feed. A ready-made map of where hypotheses usually hide is in the Google Ads account audit checklist.
One especially useful pattern: use experiments as a filter for Google’s recommendations. The interface suggests broadening match types, raising budget, enabling a new feature — instead of accepting or dismissing on faith, run the suggestion through a split test and get an answer on your own data. Which recommendations even deserve that treatment is covered in the article on optimization score and Google Ads recommendations.
The reverse link matters too: experiment results are inputs to budget planning. A change validated by a split test can be baked into a forecast; an unvalidated one cannot. How to extract an incremental cost per conversion from that forecast is in the piece on Performance Planner and budget planning.
Three worked scenarios
Scenario 1: the bid strategy test that “lost” and won
A campaign delivers 60 conversions a week at a $35 tCPA. Hypothesis: switching to Maximize Conversion Value with a 350% tROAS raises revenue at the same spend. Split 50/50, four weeks.
| Metric | Control (tCPA) | Variant (tROAS) | Delta |
|---|---|---|---|
| Spend | $4,200 | $4,180 | −0.5% |
| Conversions | 121 | 104 | −14% |
| Conversion value | $18,900 | $21,300 | +12.7% |
| ROAS | 450% | 510% | +13% |
By conversion count the variant lost 14%, and reflex says roll it back. By money it won: the same spend produced $2,400 more revenue. This is “winning on the wrong metric” running in reverse — the test was about value, so it has to be judged on value. The decision goes to the metric named in the hypothesis before launch, which is precisely why you write it down first.
Worth noting: for the first ten days the variant trailed on both metrics because the strategy was still learning. Stopping in week two would have produced the opposite conclusion.
Scenario 2: significant but not worth shipping
A large account, 1,400 conversions a week. A new asset framing tests at +2.1% conversions with 96% confidence. Formally a win. Now the money: +2.1% at a $22 CPA on $30,800 weekly spend is roughly 29 extra conversions a week, around $640 in incremental profit at typical margin. Rolling the new asset set across 40 campaigns costs two days of specialist time plus design. It pays back — over a month, and only if the effect holds. Verdict: ship it, but in the normal queue rather than as a priority. Significance is admission to the discussion, not the decision.
Scenario 3: a false conclusion from thin data
A campaign with 18 conversions a week. New landing page tested over three weeks: 27 conversions for control, 21 for the variant. A 22% gap looks convincing and is not significant — at those counts the confidence interval spans both −40% and +25%. The correct conclusion is that the test produced no answer. The common, incorrect one is “the new page is worse, roll it back.” That’s exactly how teams discard working hypotheses and spend years rediscovering things they already tested, because the decision rested on numbers that meant nothing.
When a split test isn’t possible
Not every hypothesis fits inside a campaign. The account may be too small, the change may touch the whole site, or you may be testing a call center process or an entirely new offer. Alternatives, in descending order of reliability:
- Geo split. Divide regions into two comparable groups and enable the change in one. Works even where in-platform experiments aren’t available, and captures effects that live outside the ad account. Requires groups matched on volume and demand structure.
- Staged rollout. Ship to 20% of campaigns, then 50%, then all, checking metrics at each step. It doesn’t establish causality, but it caps the damage from a bad decision.
- Matched campaign control group. Apply the change to half of a set of similar campaigns and leave the rest untouched. Weaker than a split (campaigns are never identical) but far better than before-and-after.
- Switchback testing. One week on, one week off, repeated four to six times. Alternation cancels seasonality. Suits fast-acting changes; unsuitable where a learning period exists.
- Plain before-and-after. The last resort. If you must use it, at minimum match period length, start on the same weekday, and write out explicitly everything else that changed in that window.
A realistic annual testing calendar
| Quarter | Testing focus | Why then |
|---|---|---|
| Q1 | Bid strategies and targets | Low season — mistakes are cheap |
| Q2 | Structure and match types | Room for a long-running test |
| Q3 | Landing pages and assets | Preparing the base for peak season |
| Q4 | Nothing structural | In peak you test only what can’t wait |
“Don’t experiment during peak season” sounds dull, but it saves more money than any single successful test: the cost of an error in November is multiples higher, and the data quality is worse because demand is abnormal.
The infrastructure precondition nobody mentions
An experiment needs an account stable enough to survive the full runtime. If the account gets suspended in week two, or spend limits prevent both groups from accumulating volume, testing becomes a lottery and you end up deciding on half a sample. For teams that need a predictable platform with real history and adequate limits, that’s an infrastructure question: agency Google Ads accounts and the rest of the PPC Rebels services exist to provide the stability that any testing methodology quietly assumes.
FAQ: experiments in Google Ads
How is an experiment different from a before-and-after comparison?
An experiment runs both versions at the same time on the same traffic, so external factors — season, competitors, your other channels — hit both groups equally and cancel out. Period comparison attributes to your change everything that happened in the world during that window.
What traffic split should I use?
50/50 by default — it maximizes sensitivity for a given volume. Give the variant a smaller share only when the change is risky, and accept that the test will take substantially longer.
How long should an experiment run?
At least two full weeks, typically two to four depending on conversion volume. Set the end date before launch and don’t move it. If a bid strategy changes inside the test, add 7–14 days for learning.
What does “95% statistical significance” mean?
It means the probability of seeing the observed difference by chance, if there were no real effect, is under 5%. It’s a risk threshold, not a guarantee — and it says nothing about the size of the effect. A significant lift can still be too small to be worth shipping.
Can I stop early if the variant is already ahead?
Technically yes, methodologically no. Stopping at the first favorable reading systematically overstates the effect because you’re catching random fluctuation. If you need a shorter test, plan for it before launch.
What should I test if I have very few conversions?
Big hypotheses with large expected effects: a different landing page, a different bid strategy, a materially different structure. Small refinements are statistically indistinguishable at low volume — implement them on judgment and evaluate the portfolio over a quarter.
What is a lift study and how does it differ from an experiment?
In a lift study a holdout group sees no ads at all, and you compare exposed against unexposed behavior. It measures the incremental effect of advertising itself, not a comparison of two ad configurations. Volume requirements are higher, and it often needs Google rep involvement.
Can I run several experiments at once?
Across different campaigns, yes. Inside one campaign, parallel tests overlap on traffic and results become unreproducible. Keep one active experiment per campaign.
What if the experiment shows no significant difference?
That is a result: the change produces no effect you can measure at your volume. Don’t ship it as an “improvement,” and don’t spend time on it again — log it and move to the next hypothesis with a larger expected effect.
Do experiments work in Performance Max?
Yes, PMax is supported in Experiment Center, notably for measuring the campaign’s contribution and for asset tests. Interpretation is harder than in search because the surface mix is opaque: you see the outcome but not which placement produced it.
How do I know a winning test should be rolled out account-wide?
Three conditions at once: the result is significant, the effect size justifies implementation, and the mechanism is understood. If the win is inexplicable, replicate it on another campaign before generalizing.