Adbrains

A/B Testing for Ads: Step-by-Step Guide to Reliable Results in 2026

Category

Google Ads

icon

Written by

Adbrains

icon

Post date

21 September 2026

Every advertiser wants to know which ad performs best. But how do you arrive at an answer you can actually trust, rather than a conclusion based on chance? A/B testing is the answer, but only when approached with structure and discipline. Too many advertisers draw conclusions from too little data, test multiple elements simultaneously, or stop a test too early because results look promising. The outcome: decisions based on noise rather than genuine performance. In this article, you will learn how to set up A/B tests for Google Ads and Meta Ads that are statistically reliable, which pitfalls to avoid, and how AdBrains AI technology radically improves this process.

What Is A/B Testing for Ads?

A/B testing, also known as split testing, is the process of simultaneously showing two or more variants of an ad element to comparable audiences, in order to measure which variant performs better on a predefined goal. That goal can be CTR, conversion rate, CPA, ROAS, or a combination of metrics. The power of A/B testing lies in isolating variables: you change only one element at a time, so you can say with confidence that any improvement found is the result of that specific change.

In Google Ads, for example, you test two headlines in a Responsive Search Ad (RSA), or compare two versions of a landing page using a Google Ads experiment. In Meta Ads, you test creative variants such as images, video thumbnails, or ad copy. Both platforms offer built-in experimentation functionality, but the way you set up and interpret the test determines whether your result is reliable.

The difference between manual testing and AI-driven testing is substantial. Where manual testing depends on discipline, spreadsheets and human decisions at exactly the right moment, an automated approach delivers consistency, speed and statistical reliability. But let us start with the foundation: the step-by-step guide to a reliable A/B test.

Step 1: Formulate a Sharp Hypothesis

A good A/B test does not start with a variant but with a question. The hypothesis is the core of every test. A strong hypothesis follows this structure: "If we change [element X] to [variant Y], we expect [metric Z] to improve, because [reason]." Without an explicit reason, you are testing blindly and cannot transfer learnings from the outcome to future campaigns.

For example: at Clima-Active.nl, an installer of air conditioning and heat pumps working with quote requests as a conversion goal, you notice that ads have a high impression share but a relatively low click-through rate. The hypothesis might be: "If we change the headline from a product name ('Daikin Heat Pump') to a problem-solving angle ('Reduce heating costs by 60%'), we expect CTR to rise, because users actively searching for energy savings respond better to a benefit than to a brand name."

Step 2: Determine Sample Size and Test Duration

One of the most common mistakes in A/B testing is ending a test too early. After three days you see that variant B outperforms variant A by 40%, and you declare variant B the winner. But that 40% could be pure chance if the total number of clicks or conversions is still too low. Statistically, you need a minimum sample size to consider a difference significant.

The rule of thumb is that you want to reach a significance level of at least 95% before drawing a conclusion. This means the probability that the observed difference is due to chance is less than 5%. To achieve this, you need a certain number of conversions or clicks per variant, depending on the magnitude of the expected difference (the so-called minimum detectable effect, or MDE). The smaller the expected difference, the more data you need.

For e-commerce campaigns like those of ToetsJeKennis.nl (online exams and courses with an average order value of around 50 euros), it is realistic to run a test for at least two to four weeks. For campaigns with higher conversion volumes you can reach significance faster. For lead generation campaigns with lower conversion volumes, like LeroyBrouwer.nl, a reliable test may take four to six weeks. Plan this in advance so you are not tempted to stop too early.

  • Use a sample size calculator to determine the minimum test duration before launch
  • Set a fixed end date and avoid checking results in between to prevent confirmation bias
  • Ensure each variant receives at least 100 conversions for reliable results in conversion-focused tests
  • Account for weekday differences: always run tests for a complete number of weeks
  • Avoid testing during exceptional periods such as major promotions or public holidays, unless that period is specifically what you want to test

Step 3: Isolate One Variable Per Test

The golden rule in A/B testing: never change more than one element at a time. This sounds simple, but in practice advertisers still make the mistake of changing both the headline, the description, and the image in a single test setup. If that test then shows a significant difference, you cannot tell which element caused it. You cannot reproduce or transfer the result to other ads.

In Google Ads RSA tests, use the pin function to lock specific headlines to fixed positions, so you know exactly which headline is being shown and keep the comparison clean. In Meta Ads, use the built-in A/B test functionality in Ads Manager to test creative variants, audiences or placements separately. Multivariate testing (testing multiple combinations simultaneously) is an advanced technique that only makes sense when you have enough volume to reach significance for every combination.

Step 4: Set Up Conversion Tracking Correctly

A test is only as reliable as your conversion tracking. If you are not tracking conversions completely and correctly, you are measuring the wrong thing, or worse: missing conversions entirely. For Google Ads, it is essential to use server-side tracking via a server-side Google Tag Manager (sGTM) setup. This prevents browser-based limitations (such as ad blockers or iOS privacy restrictions) from distorting your conversion data.

Enhanced Conversions add an extra layer of reliability by sending pseudonymised customer data (such as email addresses) back to Google, so that conversions that would otherwise be lost are still measured. When running an A/B test, you need to be certain that both variants are measured equally, without data loss in either branch. Inconsistent measurement is one of the most underestimated causes of misleading test results.

Step 5: Analyse Results and Draw the Right Conclusions

Once your test has reached the planned duration and the 95% significance threshold has been crossed, it is time to analyse the results. Do not only look at the primary metric, but also at secondary metrics. A headline that increases CTR by 34% but halves the conversion rate is not a winner. You want to see a holistic improvement: more clicks from the right audience, who also actually convert.

Always document the result of every test, even when no significant difference is found. A null result is also a result: it tells you that the tested variable has less impact than thought, and helps you prioritise future tests. Build a test library per campaign and per client, so knowledge accumulates and you do not have to repeat the same tests.

  • Look at the 95% confidence interval, not just the point estimate
  • Break down results by device (mobile vs. desktop) if the volume allows it
  • Check whether external factors (season, news, competitor changes) influenced the test period
  • Compare the winning variant against the historical baseline from before the test
  • Immediately plan the follow-up test based on insights from the current test

Overview: Elements You Can A/B Test Per Platform

Element Google Ads Meta Ads Primary Metric
Headline RSA pin test, campaign experiment Primary text variation CTR
Description text RSA variants Body text variation CTR, Ad Strength
Image / creative Display and PMax asset test Creative A/B test CTR, ROAS
Call-to-action RSA variants, sitelinks CTA button variation CTR, conversion rate
Landing page Campaign experiment (URL split) URL variation per ad set Conversion rate, CPA
Bidding strategy Smart Bidding experiment (tCPA vs. tROAS) Not directly applicable CPA, ROAS
Audience Audience experiment via campaign split Audience A/B test CPA, CPL

The table illustrates how broad the applications of A/B testing are. Whether you want to compare a Smart Bidding strategy against an alternative, or test two creative angles for a Performance Max campaign, in all cases the keys to reliable insights are the same: structured isolation of variables and sufficient data.

How AdBrains Automates A/B Testing with Its Own AI Technology

At AdBrains, we have organised A/B testing not as a periodic activity, but as a continuously automated process that runs in parallel with daily campaign optimisation for every client. Our own AI technology takes over the entire testing process, from hypothesis formation to rollout of the winner, and does so faster, more consistently and more reliably than manual management ever could.

Our RSA improvement system continuously monitors the Ad Strength scores of all ads. As soon as an ad has a POOR or AVERAGE score, the AI automatically analyses which element is the weakest link: a generic headline, a description that is too short, or a missing call-to-action. Based on that analysis, the system generates alternative variants that are set up as a controlled A/B test, with automatic calculation of the required duration based on the historical conversion volume of that specific campaign.

Our multi-agent verification system plays a crucial role here: four independent AI agents evaluate every proposed test setup before it goes live. They verify that only one variable has been changed, that the sample size is achievable within the available campaign period, and that the primary metric matches the client objective. This prevents the most common errors that manual testers make.

After reaching statistical significance (we standardly use 95%), the system automatically applies the winning variant as the active ad. The losing variant is archived in our test library with all associated metadata: duration, volume, significance level and the tested hypothesis. This cumulative learning process ensures that our AI becomes increasingly accurate at predicting which variants are promising for clients like ToetsJeKennis.nl and E-4motion.com, making test cycles progressively more efficient.

Additionally, our server-side signal enrichment system links test results directly to first-party conversion data. This ensures both test variants are always measured under identical, uncontaminated conditions, even when browser-based tracking partially fails. This is critical for reliable conclusions, especially for campaigns at clients like Clima-Active.nl where every quote request carries significant value and data loss directly undermines test integrity.

Frequently Asked Questions About A/B Testing for Ads

How long should an A/B test run to produce reliable results?

The minimum duration depends on your conversion volume and the expected difference between variants. As a rule of thumb: run a test for at least two complete weeks (to neutralise weekday differences) and ensure at least 100 conversions per variant for conversion-focused tests. For CTR tests on high-volume campaigns you can reach significance faster, but always use a minimum of seven days. Use a sample size calculator to determine the exact duration based on your baseline conversion rate and desired minimum detectable effect.

Can I run multiple A/B tests simultaneously on the same campaign?

No, not on the same campaign. Parallel tests on the same campaign influence each other's results because they compete for the same budget and audience. Use Google Ads experiments to isolate tests cleanly: an experiment splits traffic at the campaign level and ensures both variants get equal opportunities. For multiple simultaneous tests, use separate campaigns with individual budgets and audiences.

What is the difference between A/B testing and multivariate testing?

In A/B testing you compare two variants of one element (e.g. headline A vs. headline B). In multivariate testing you test multiple combinations of multiple elements simultaneously (e.g. headline A and B, combined with description X and Y). Multivariate testing requires significantly more traffic to reach significance for each combination. For most advertisers, A/B testing is the practical choice; multivariate testing only makes sense for campaigns with very high traffic volumes.

Which metric should I choose as the primary metric for my A/B test?

Always choose the metric that most closely aligns with your ultimate campaign goal. For brand awareness campaigns, that is CTR or impression reach. For Performance Max and Search campaigns focused on e-commerce, the ROAS or conversion rate is the right primary metric. For lead generation campaigns, such as those for LeroyBrouwer.nl or Clima-Active.nl, the Cost Per Lead (CPL) or the conversion rate on the request form is the best choice. Avoid using a proxy metric (such as CTR) as your primary metric when you actually have a conversion goal, unless you specifically want to measure only the click attractiveness of the ad.

How do I know if my A/B test result is statistically significant?

Statistical significance means that the observed difference between two variants is, with a certain degree of certainty, not due to chance. The industry standard is a significance level of 95%, meaning the probability of a false positive result (a Type I error) is less than 5%. You calculate this with a z-test or chi-square test, or use an online A/B test calculator. Many advertising platforms, such as Google Ads in the experiments module, automatically calculate statistical significance and provide a recommendation when the test is ready to conclude.

Share this article

Let a Google Ads Expert review your current campaigns

In a personal call we analyze your current Google Ads setup and show concrete improvements. Free and non-binding.

Account Analysis

Within 30 minutes

We dive live into your Google Ads account and pinpoint quick wins for a higher ROAS.

AI Platform Demo

Live walkthrough

See how our AI analyzes search terms daily, optimizes bids and expands your campaigns.

Tailored Growth Plan

Concrete action plan

You get a clear plan with expected results, a timeline and investment for your webshop.