...
Google Ads Tutorials

Google Ads Experiments: The Right Way to A/B Test (Without Destroying Performance)

Dan Kabakov, Google Ads Certified Partner Last Updated: April 1, 2026 5 min read

Use Google Ads Experiments for A/B testing — not duplicate campaigns. Experiments use controlled traffic splits that prevent auction competition, giving you clean data without destroying performance in either variation. The most valuable test for e-commerce: Performance Max vs Standard Shopping.

You want to test whether Performance Max outperforms Shopping campaigns. Or whether broad match keywords beat exact match. Or whether new ad assets improve conversion rates.

The old approach: run two campaigns simultaneously and compare results. The problem: those campaigns compete against each other in the same auctions, driving up your costs and making both perform worse. You end up with garbage data and wasted budget.

The solution: Google Ads Experiments. Experiments let you properly A/B test different campaign variations with controlled traffic splits. No competition, no chaos, no inflated costs. Just clean data showing which approach actually works better.

In this guide, I'll show you exactly how to set up experiments, which tests are worth running, and which "experiments" Google pushes that you should probably avoid.

01Why A/B Testing in Google Ads Requires Experiments

Testing is crucial in digital advertising. Without it, you're guessing which audiences, creatives, campaign types, and settings perform best.

But here's what most advertisers get wrong: they try to A/B test by running duplicate campaigns.

The Old (Wrong) Way

  1. Create Campaign A (original)
  2. Create Campaign B (variation)
  3. Run both simultaneously
  4. Compare results after 30 days

Why This Fails

  1. Campaigns compete in auctions — you bid against yourself, inflating CPCs for both
  2. Budget splits unevenly — Google favors one campaign, starving the other of data
  3. Attribution confusion — same users see both campaigns, muddying conversion data
  4. Performance tanks — both campaigns perform worse than either would alone

This is especially disastrous with Performance Max. Put two PMAX campaigns targeting the same products, and they'll cannibalize each other into oblivion.

The Right Way: Google Ads Experiments

Experiments solve this by using controlled traffic splits — each variation gets an exact percentage of traffic. They don't bid against each other, users are assigned to one variation only, and Google calculates confidence levels for results. Same test, clean data, no performance destruction.

02How to Access Google Ads Experiments

The Experiments feature is somewhat hidden. Here's how to find it:

1
Open your Google Ads account

Log in at ads.google.com and select the account you want to test in.

2
Click "Campaigns" in the left sidebar

Expand the Campaigns section to reveal the submenu.

3
Click "Campaigns" again in the submenu

The submenu includes Campaigns, Ad groups, Ads, and Experiments.

4
Select "Experiments"

This opens the Experiments dashboard where all active and past experiments are listed.

5
Click "+ Create Experiment"

You'll see experiment type options — choose the one that fits your test hypothesis.

03Types of Google Ads Experiments Available

When you create an experiment, you'll see several options. Not all are equally valuable. Here's a breakdown of each one and whether it's actually worth running.

1. Campaign Features and Settings

Tests variables like keywords, conversion goals, and settings for a single campaign. Three sub-options exist:

Final URL Expansion in Performance Max

Tests automatically created assets and lets Google AI send traffic to relevant landing pages and generate text assets to match search queries.

My take: Skip this. Giving Google full control to automatically create assets can backfire, especially for established brands. Google might generate messaging you don't want associated with your brand. Only consider it if you're a small business with limited creative resources and no brand guidelines to protect.

Broad Match Keywords for Search

Tests broad match versus your current match types (phrase or exact). Google pushes this aggressively because broad match means you spend more. It CAN discover valuable long-tail queries, but it also wastes significant budget on irrelevant searches. If you want to test this, run it for at least 4 to 6 weeks with strict monitoring. Don't enable it and walk away.

AI Max

Tests showing your ads in Google's AI mode search results using broad match keywords. Want to advertise in AI search results? Fine — but Google requires broad match in return. Worth testing if AI search is growing in your industry, but monitor spend carefully.

2. Assets

Tests different asset variations — images, videos, callouts, sitelinks. When you're running Responsive Search Ads, they already have 10 to 15 headline variations testing automatically. Asset experiments add a redundant layer on top of that. Use one well-constructed RSA per ad group and monitor performance reports to see which headlines perform best.

3. Campaign Types

Tests different campaign types against each other. This is where experiments become genuinely valuable for e-commerce advertisers. Available comparisons: Performance Max vs Search, Performance Max vs Display, Performance Max vs Shopping.

Performance Max vs Shopping (Most Useful)

The most common and valuable test for e-commerce: does PMAX or standard Shopping perform better for your products? This experiment gives you clean data to answer that question without the campaigns competing against each other. Highly recommended for any e-commerce store debating between these two campaign types.

Performance Max Uplift (Be Careful)

Tests the impact of adding Performance Max to your existing campaigns. This experiment has a major flaw: Google's data-driven attribution model gives PMAX credit for conversions that other campaigns actually drove. PMAX prioritizes bottom-of-funnel users who were going to convert anyway, then claims the credit. Don't rely on this experiment to prove PMAX value. For accurate PMAX analysis, use third-party tools like Northbeam.

4. Custom Experiments

Select a campaign type and create your own experimental variation — maximum flexibility for specific hypotheses you want to test. This is the right choice when none of the prebuilt types match your question.

Google Ads Masterclass
Master Testing and Optimization in Google Ads
Learn advanced testing strategies, campaign structure frameworks, and optimization techniques built for e-commerce stores. Includes real account walkthroughs and the exact frameworks I use with clients.
Explore the Course

04How to Set Up a Shopping vs Performance Max Experiment (Step-by-Step)

Let me walk through setting up the most valuable e-commerce experiment: testing whether PMAX or Shopping performs better for your products.

1
Select Experiment Type

Go to Experiments dashboard, click "+", select "Campaign Types", then choose "Performance Max vs Search, Display, or Shopping". Click Continue.

2
Select Control Campaign

Choose "Shopping" as your campaign type, then select your existing Shopping campaign as the Control. Important: choose a campaign that is not already in an experiment.

3
Configure Traffic Split

Google defaults to a 50/50 split — Control (Shopping) vs Treatment (PMAX). This ensures no auction competition and a fair comparison. You can adjust to 70/30 if you want to minimize risk, but that extends the time needed to reach statistical significance.

4
Choose Success Metric

For e-commerce, select Conversion Value (measures revenue). For lead generation, select Conversion Volume. You can also set a specific target ROAS — the experiment will show which campaign type better achieves your goal.

5
Set Duration

Google suggests 3 months for PMAX experiments, which is reasonable — PMAX needs 2 to 4 weeks to exit learning phase, plus additional weeks for optimization and statistical significance. Minimum: 6 to 8 weeks. Ideal: 12 weeks. Don't cut experiments short.

6
Name, Schedule, and Launch

Give the experiment a clear name (e.g., "PMAX vs Shopping Q2 2026"), confirm the political ads setting (No for most advertisers), and click Schedule. The experiment is now live — monitor it from the Experiments menu.

05Reading Experiment Results: What to Look For

After your experiment runs for sufficient time, here's how to interpret the results properly.

Key Metrics to Compare

Metric What It Tells You Priority
Conversion Value Which variation generated more revenue Primary
ROAS Return on ad spend — efficiency metric Primary
Conversion Volume Total number of sales (useful secondary check) Secondary
Cost per Conversion Which variation was more efficient per sale Secondary
Budget Delivery Did both variations spend their allocated budget? Secondary

Confidence Level

Google calculates statistical confidence for experiment results. Here's what to do with each level:

  • 95%+ confidence — results are statistically significant, safe to act on
  • 80 to 95% confidence — trending positive but might need more time to confirm
  • Below 80% — not enough data to conclude, extend the experiment
Attribution Warning

When comparing PMAX to other campaign types, remember: Google's data-driven attribution model favors PMAX. If PMAX shows 20% better results, the real difference might be 5 to 10%. For accurate PMAX evaluation, supplement Google's experiment data with a third-party attribution tool before making major budget decisions.

06Which Experiments Are Actually Worth Running?

Based on real-world results, here's how to prioritize your testing roadmap.

High Value

  1. PMAX vs Shopping — essential for e-commerce stores unsure which approach works better for their catalog
  2. Bidding strategy tests — compare Target ROAS vs Maximize Conversion Value on the same campaign
  3. Audience signal tests — different audience segments in PMAX asset groups
  4. Landing page tests — different final URLs for the same products

Medium Value

  1. Broad match keywords — worth testing if you have budget for discovery and can monitor it closely
  2. AI Max — test if AI search volume is growing meaningfully in your niche
  3. Custom campaign variations — for specific hypotheses about individual settings

Skip or Approach with Caution

  1. Final URL Expansion — too much brand control risk, Google generates your messaging
  2. PMAX Uplift — attribution issues make results structurally unreliable
  3. RSA ad variations — RSAs already test internally across headlines
  4. Auto-created assets — quality control is a genuine concern

07Common Experiment Mistakes to Avoid

Mistake 1: Ending Experiments Too Early

After two weeks, one variation looks clearly better and you want to call it. The problem: two weeks isn't enough data. Early trends often reverse, and learning phases skew initial results heavily. The rule: minimum 6 weeks for any experiment, 12 weeks for PMAX.

Mistake 2: Testing Too Many Variables

You change bidding strategy AND targeting AND creatives in your experimental variation. If results differ, you won't know which change caused it. The rule: one variable per experiment. Test bidding strategy first, then targeting, then creatives — separately.

Mistake 3: Uneven Budget Allocation

You give the experimental variation 20% traffic to minimize risk. The problem: 20% traffic takes 5x longer to reach statistical significance. The experiment runs for months and may never conclude. The rule: use a 50/50 split for fastest, most reliable results. Only use uneven splits if you genuinely cannot risk equal budget allocation.

Mistake 4: Ignoring External Factors

You run an experiment during a major sale or seasonal spike. One variation "wins" by 40%. But seasonal spikes don't represent normal performance — results often reverse during regular periods. Run experiments during normal business periods. If you must test during seasonal peaks, extend the experiment to include non-peak weeks.

Mistake 5: Trusting Google's Attribution Blindly

Experiment shows PMAX crushed Shopping. You move all budget to PMAX. Overall performance drops. PMAX attribution is inflated — it took credit for conversions other campaigns would have captured. For PMAX experiments, verify results with third-party attribution data before making major budget shifts.

08Best Practices for Google Ads Experiments

Before Starting

  • Document current performance benchmarks so you have a clear baseline
  • Define clear success metrics before you begin — decide what "better" means upfront
  • Calculate minimum sample size needed for statistical significance
  • Plan experiment duration — minimum 6 weeks, 12 for PMAX
  • Identify external factors that could affect results (sales, seasonality, competitors)

During the Experiment

  • Do not make other changes to the campaigns being tested
  • Monitor weekly but do not react prematurely to early data
  • Check for technical issues — ads disapproved, budget caps hit, learning phase stuck
  • Document any external events that could influence results

After the Experiment

  • Wait for 95%+ confidence before declaring a winner
  • Verify results with secondary metrics, not just the primary KPI
  • Consider attribution model impact — especially for any test involving PMAX
  • Implement the winner gradually rather than switching 100% of budget immediately
  • Document learnings for future tests — every experiment should build on the last
Free Resource
Get the Free Google Ads Audit Template
Before running experiments, make sure your account fundamentals are solid. This 47-point audit template reveals conversion tracking issues, campaign structure problems, and optimization opportunities you can fix before testing anything.
Get the Free Template

09The Bottom Line: Test Properly or Don't Test at All

A/B testing is essential for Google Ads optimization. But bad testing is worse than no testing — it gives you false confidence in wrong conclusions.

The wrong way: running duplicate campaigns that compete against each other, comparing results, and calling it "testing."

The right way: using Google Ads Experiments with controlled traffic splits, sufficient duration, and clear success metrics.

What to Test

  • PMAX vs Shopping (if you're e-commerce — this is the big one)
  • Bidding strategies: Target ROAS vs Maximize Conversion Value
  • Audience segments in PMAX asset groups
  • Landing page variations using different final URLs

What to Approach Carefully

  • Broad match keywords — Google wants you to spend more, so monitor the data honestly
  • AI Max features — the inventory expansion comes with trade-offs
  • Any PMAX-involving attribution — results are structurally inflated

What to Skip

  • Auto-created assets — brand risk is real and hard to reverse
  • PMAX uplift tests — attribution problems make results unreliable
  • Redundant RSA variations — RSAs already handle this internally

Experiments are powerful tools when used correctly. Set them up properly, let them run long enough, interpret results carefully, and you'll make data-driven decisions that actually improve performance. For more on the campaign types you're testing, see the complete guide to Performance Max optimization and the portfolio bidding strategy guide.

10Frequently Asked Questions

Run experiments for a minimum of 6 to 8 weeks. For Performance Max experiments, 12 weeks (3 months) is recommended because PMAX needs 2 to 4 weeks to exit the learning phase, plus additional time for optimization and statistical significance. Cutting an experiment short because early results look promising is one of the most common mistakes — early trends often reverse.
Running duplicate campaigns creates auction competition where you bid against yourself, inflating CPCs for both. Budget splits unevenly because Google's algorithm favors one campaign, starving the other of impressions and data. Attribution gets muddied when the same users see both campaigns. Both campaigns perform worse than either would alone. Google Ads Experiments use controlled traffic splits that prevent this competition entirely — users are assigned to one variation only.
Yes — this is one of the most valuable experiments for e-commerce stores. The experiment gives you clean data comparing the two campaign types without them competing against each other. However, be aware that Google's data-driven attribution model tends to favor PMAX, which inflates its apparent performance. Verify results with third-party attribution tools before making major budget shifts based on experiment data alone.
Use a 50/50 split for the fastest and most reliable results. Uneven splits like 80/20 take significantly longer to reach statistical significance — a 20% experimental variation needs 5x more time to accumulate enough data. Only use an uneven split if you genuinely cannot risk equal budget allocation, and plan for the experiment to run correspondingly longer.