• Skip to main content
  • Skip to primary sidebar
  • Home
  • About PPC Chat
    • Where to Find PPC Chat
    • The History Of PPC Chat
    • Who Is Behind PPC Chat?
    • PPC Chat Badge
  • Resources
    • Suggest A Topic
    • Ad Platform Resources
    • Ad Platform Help Centers
    • PPC Acronyms
    • Google Analytics 4 (GA4) Resources
    • Privacy & Cookieless Resources
    • Podcasts You Might Like
    • Experienced Speakers Helping New Speakers Program
  • Support PPC Chat
    • PPC Chat Swag
  • Sponsors
  • Raise Someone’s PPC Profile
    • Recognize A PPCer
    • PPC Profile Spotlights
    • New Speaker Support Program

PPCChat

The Official Home of the Twitter Chat

How To Design Effective A/B Tests for PPC

Guest post by Reid Thomas

One of the key differentiators cited to use digital marketing rather than traditional is that it’s easily measurable. The numbers and metrics are the thing for so many PPC folks – managing those numbers is a huge part of the job!

Here’s the problem: we can collect a lot of data, and easily display it in tables and graphs and reports. It’s easy to see that one number is bigger than another, and know to reward the better-performing ad, campaign, or network.

But how do we know – really know – if one campaign is better than another? How can we put rigor to our tests to make sure that we’re really making sure that we’re overcoming the statistical challenges that come with measuring so much data? How do we design effective tests that help us make the best decisions for our clients, bosses, and stakeholders?

Testing Like It’s 1923

Like so many things in digital advertising, it’s not totally true that measurability started because of computers. Some might disagree about its practical use a century later, but Claude Hopkins lays out the case for using what he calls “keyed advertising” in his century-old book, Scientific Advertising.

From Chapter 1:

One ad is compared with another, one method with another. Headlines, settings, sizes, arguments and pictures are compared. To reduce the cost of results even one per cent means much in some mail order advertising.

And again from Chapter 15:

Almost any questions can be answered cheaply, quickly and finally by a test campaign.

If this sounds familiar, then welcome to testing like it’s 1923! Hopkins’s work came up alongside modern statistical analysis around the turn of the century. It’s unknown whether or not he was particularly aware of or interested in becoming more statistically robust in his practice. It was enough for him to act as many of us do today: picking winners by which test “won.”

But to move perhaps towards the latter half of the 20th century, we need to consider three topics for each test we want to run:

  • what sample size we use,
  • what our tested hypothesis is,
  • and how we report on our test.

By taking into account these topics, we can really answer “almost any question” “cheaply, quickly, and finally.”

Sample Sizes & Statistical Significance

One of the most troubling questions we get as advertisers running a test is, “Are you sure?”

It’s an important one! The people making the decisions about their budgets deserve to be sure about how their money’s being managed. And you as the advertiser should be able to say with confidence that the data means what it says.

We have 4 possible outcomes of any test we set up:

  1. the data says the hypothesis isn’t true, and the reality is that the hypothesis isn’t true – a true negative;
  2. the data says the hypothesis isn’t true, but actually it is true and we got a false negative – a “type II error” if you’re fancy;
  3. the data says the hypothesis is true, and the real world reflects the data – a true positive;
  4. and the most dangerous one, that the data says the hypothesis is true, but the reality is that data is a false positive – also known as a “type I error.”

Our goal as advertisers is to make the correct choice, not just based on the data we’ve observed, but in reality. This is our main goal with reaching what’s called “statistical significance.” We want to minimize the chance that the data we have in front of us just happened due to random chance, rather than because of our hypothesis.

Confidence and Power Levels

There’s a ton of statistical methods to minimize the chance for type I and type II errors if we were scientists. One of the most common is to run our test enough times to reach a given “p-value” or “confidence interval.”

We often see stuff like “95% confidence interval” or “5% error” when we look at studies and tests. This means that we’ve run a test with enough samples to assume that there’s only a 5-in-100 chance that a positive test result happened because of something other than what we’re testing. Put another way, one 1 of every 20 positive test results will be a type I error.

On the other hand, we also want to avoid type II errors. We can do this by ensuring a high enough “power level.” Just like we typically use a 95% confidence interval, we can use convention to help us choose a power level. We typically use a lower power level – 80% – because we are less concerned if we have a false negative.

It’s important to keep these chances in mind. Is it acceptable to us, our team, and our stakeholders that 1 of every 20 winners is a type I error? What about that 1 out of 5 losers is actually a false negative? Our risk level might depend on many factors, but we can tighten or loosen the confidence interval and power level with intent to best match our risk.

Reaching Statistical Significance

We use the confidence and power level to give us two of the important numbers we need to figure out how many times a test needs to be run for us to be confident in our test – the appropriate sample size.

Using a tool such as the sample size calculator from AB Testguide, we can quickly figure out how many impressions, clicks, or visits we need to make a valid test.

Using this calculator, we need the confidence and power levels that we determined from before, but there’s a few other points of data to identify our statistically significant sample size needed:

  • our current conversion rate, in percentage;
  • our expected improvement, in percent change of percentages – i.e. a 10% difference in conversion rate would be the difference between 11% and 10%;
  • whether our hypothesis is one- or two-sided – i.e. whether we want to hypothesize that the variant is greater than the original or simply different than the original.

This calculator also uses the unique visitors on your test page per week and the maximum number of weeks you have to test to give you additional information.

Wow, That’s a Big Sample Size

It’s really common to see surprisingly large sample sizes for tests.

For example, let’s suppose a test like this for a lead generating landing page:

  • 95% confidence and 80% power, using standard benchmarks;
  • a current conversion rate of 8%;
  • expecting an 11% increase – that is, we expect the test should increase to 8.86%;
  • using a one-sided test to see if the variant is “better” than the original only.

The sample size for such a test is 12,338 visitors per variant! With $8 per click and 4000 total visitors per month, we’d be looking at over $200,000 spent over 6 months to reach statistical significance!

The benefit of using a calculator is that we can rapidly figure out the numerical side of our test: what do we need to achieve to be able to run a test in a reasonable amount of time and money? How much of a change do we need to see? What kind of tests can we run on the campaigns as they’re running? These questions help us determine our hypothesis.

Designing a Testable Hypothesis

By looking at the metrics that we’ve measured in the campaign, we can start to develop a full hypothesis for testing. There are various frameworks that we can use, but for business-focused marketers, the SMART criteria is a touchpoint that crosses departments and functions.

SMART gives us five criteria for our hypothesis:

  • specific, forcing us to choose our metrics so that we can calculate an effective sample size;
  • measurable, an inherent part of testing on metrics;
  • actionable, ensure that we tie the hypothesis to a larger trend or observation so that it’s more broadly applicable across our campaigns and accounts;
  • realistic, tied to a percentage increase that we feel provides confidence in our results while achieving a result in a reasonable time-frame; and
  • time-bound, reminding us to reject the hypothesis if the test doesn’t reach a conclusion by the end of the test.

A hypothesis built this way might read:

Because of better visibility on the page, we expect that increasing the font size of the CTA link by 1rem will increase the CTR of that link by 50% from 1.5% to 2.25% over the course of four weeks.

If the page gets 2000 visitors a week, we can run this test and get a result in a reasonable length of time.

The Importance of Insight

Potentially the most important addition the SMART framework gives us is an “Actionable” hypothesis. We want to test things not only to get our individual website or client working slightly better, but so that we can help all our efforts.

To make sure we can generalize our test, we have to do more than what’s often suggested for A/B tests – we have to do more than test things at random.

If we are testing headlines, we should identify why the headlines may not be resonating and tie the changes back to the audience’s perception. We shouldn’t just say “let’s test some new creative because it’s getting stale;” we should say “let’s test some new creative because this creative is focused on technical features but doesn’t highlight the situations that our personas use our product in.”

This restatement of the goal of the test allows us to test similar ideas in other campaigns – even at the same time – to overcome sample size issues or create new best practices for organizations. This is what creates the value in a testing-focused culture: the accumulated knowledge of multiple tests in multiple situations.

Choosing the Right Metrics

The goal is to refine that insight into a measurable test. We need to be able to answer the question: how do we tell whether our generalized insight is true? We do this by figuring out two numbers: the member and the action.

The member is how we tell the test has been run once. It could be an impression of an ad or a visitor to a site. On the other hand, it could be an email sent or an ad clicked. Like with any data analytics project, it’s critical to talk about the same kinds of data so that we are specific in our reporting.

One issue with determining the member metric is the scope at which the test is run. Is it an impression-based metric, where one user might see both variants over multiple interactions? Or is it a user-based metric, where we use cookies and other methods to ensure that once a user is in a test group, they don’t see the other variant?

The action is easier to determine: what are we measuring as a successful user interaction? That said, it’s important to measure the right action, making sure it is measurable and realistic. We can choose something like a conversion, a click, or some other kind of event that represents the insight’s conclusion to measure.

Finally, we should understand and note the interaction of the action metric and the member metric. Because member metrics are either user- or impression-based, we can have very different conclusions based on supporting metrics. As much as possible, our goal should be to be aware of issues.

Once we’ve made this choice, we’d enter that value as the Conversions in the calculator.

Interpreting and Reporting on Test Results

So we’ve created an effective A/B test! We have a hypothesis that is based on a specific and actionable insight.

We’ve determined measurable and realistic metrics. And we’ve set the expectation that the effects on those metrics will reach statistical significance if the test is successful. Finally, we’ve kept the test within the time we projected initially for the campaign to reach significance.

The exact method to implement an effective test generally doesn’t matter. Whether we use a hand-coded cookie that interacts with the GTM Data Layer or use systems like a landing page’s built-in A/B test tools, our hypothesis, sample size, and test design wouldn’t change. However we implement it, we can be sure that our test will have much greater weight and value to the campaign and our organization than if our test wasn’t grounded in a strong hypothesis and clear statistical backing.

When evaluating the test, we should always validate the assessments of the tools we use. Because we thought through our test design, we may come to a different conclusion than our tools or be using a different set of metrics for our tools. For example, we could use the frequentist test calculator from AB Testguide. In that case, we’d put the member metric as the value in the “Visitors” field and the action metric in the “Conversions” field.

The Big Moment

There’s a bit of trepidation every time any of us hits the “Calculate” button. Were we right? Did the test succeed? Will it be an easy conversation or a hard one with the stakeholders in the campaign?

The most important thing with any outcome is to not fudge the numbers. We set our test up at the beginning with the knowledge that it’s a test and might fail. It’s in everyone’s best interest to take the result at face value and iterate on hypotheses from that point forward.

The second most important thing is to avoid “just doing it anyway.” We’ve all been guilty of seeing the results as inconclusive but “knowing” that the directional data is enough. We use statistical tests to ensure that we’re testing in a more rigorous way than our colleagues from a century ago and truly be data-driven advertisers, not just data-collecting advertisers.

Telling Others

At the end of the test, we need to discuss the results, even if it’s just internally. Because we did this rigorously, we can use a framework for our reporting that explains the process, our conclusions, and our next steps.

First, we discuss the hypothesis. What data gave us the hint that something could be improved? What insights drove our choices for what elements of our marketing to test? What metrics did we decide on for the test? What elements, metrics, and insights did we reject or side-table? When we discuss our thought process, we provide the context for why we would test what we decided to test.

Then, we discuss the numbers. It’s critical that we talk about the statistical significance of the test, both to explain our conclusions but also to provide a reminder that we’re doing this in a methodologically sound way. If we’re publishing the test, providing all of our methodology – what tests we used, what calculators we used, and ideally the dataset can ensure that our data can help the wider community as well.

Finally, we discuss what we’re going to do because of the test. When a test is successful, the next step is clear: implement the test variant for all users and validate that we’re seeing similar increases. But tests that fail may be more interesting to report on. What assumptions do we need to change? Is there a different concern that is overriding our insight from the hypothesis? What do we test next?

By talking about all three stages: the sample, the hypothesis, and the next steps, we can report on the test in a reasoned and effective manner.

Diving Deeper

Making effective A/B tests from design to reporting ensures that we do our best work as data-driven marketers. We can stand on the shoulders of best practices from 100 years ago to drive real impact in our organizations and help communicate with and educate our stakeholders about the audiences we’re trying to reach. But that doesn’t mean there’s no room to grow our understanding of how to do A/B tests better.

Frequentist vs. Bayesian

The statistics we learn early on are what’s called “frequentist” statistics. A/B tests largely use what’s called z-scores to determine statistical significance. But especially when working with small sample sizes, it can be useful to understand the numbers a bit differently.

Frequentist statistics wants an answer: does this test have a 19-in-20 chance of not being a false positive?

On the other hand, Bayesian statistics asks: what is the probability that the variation outperforms the original?

This changes the overall conclusion significantly. Let’s take this example:

  • Original: 2000 visitors, 200 conversions; vs.
  • Variant: 2000 visitors, 205 conversions.

The standard A/B test would see a one-sided score with a p-value of 0.3966 – far lower than the 0.95 we’re typically looking for. We’d reject this test and move on.

On the other hand, a Bayesian approach reads this as a 60.5% chance that the variant is the better choice. With a pre-set risk assessment, we can make a choice in implementation that represents our risk profile more accurately.

The challenge is that Bayesian methods use math that is far more complicated than frequentist statistics. It’s a method that less folks will understand out of the gate. And the question of setting a pre-set risk assessment is prone to guesswork at best until we’re comfortable working in that mindset.

That said, it’s an option, especially for significantly smaller sample sizes than what might get a significant result using a traditional frequentist test.

Other Ways to Measure Variance

The largest challenge of the testing used within the standard A/B test is that it assumes a far-too-simple distribution. The testing in a traditional A/B test uses a normal curve – a simple single-peak, curving mountain – to represent the range of possible outcomes.

That said, there hasn’t been testing to prove that assumption. Based on some folks’ experience, it may be a simple sloping curve where the majority of the curve under the distribution would be close to zero, and it curves downward to a 100% conversion rate. There are different tests to run, such as F-scores, that follow this distribution instead.

However, again, we’re talking above the typical pay grade of an advertiser. While it’s interesting to consider different underlying assumptions, larger organizations that are dedicated to digital marketing data science can provide better guidance.

A/A Testing to Stress Test Implementations

A final extension of this approach is to run what’s called an A/A test. Typically in a test, we’d test two variants: the original and a new variant. In an A/A test, we’d test two separate groups that receive the same original variant.

While this sounds like a redundant effort, this is one of the most valuable tools in a data-driven advertiser’s arsenal. We can ask ourselves: what’s the difference between two samples? If we use our implementation but get statistically significant results, we may have issues in how we implement our tests. By testing our tools against what should be an acceptance of the null hypothesis – that we should not see a difference because there is no difference – we can ensure that our A/B tests are actually valid.

Putting It All Together

It’s deeply important to being a modern, data-driven advertiser to take an effective approach to A/B testing.

We need to create hypotheses that are rooted in qualitative insights. We need to create tests that have metrics that we can justify. We need to report on those metrics in a way that ensures that our methodology, testing strategy, and production steps are valid and reflect our goals. And we need to go far beyond the capabilities of a century ago to generate real business results for our clients, stakeholders, and bosses.

But by thinking about all these things, we can elevate our testing beyond how our colleagues would have tested over a century ago.


Reid Thomas is an ethical digital marketer with 15 years of experience working in-house and at agencies to drive audience-centered advertising that leads to real business impact. They work through a lens of inclusivity, accessibility, brand safety, and customer privacy at their consultancy Magniventris.

Filed Under: Guest Post, Testing

Primary Sidebar

Suggest A Topic

Have an idea for a chat topic?

Submit it here!

Follow #ppcchat On Twitter

RickDronkers avatar Rick Dronkers @RickDronkers ·
11h 2099117118390223074

Hey #PPCchat whats the deal with minimum send costs being mandatory in Google Ads even for single super cheap items? Google enforcement does not seem to be consistent

Reply on Twitter 2099117118390223074 Retweet on Twitter 2099117118390223074 0 Like on Twitter 2099117118390223074 0 Twitter 2099117118390223074
RickDronkers avatar Rick Dronkers @RickDronkers ·
11h 2099117118390223074

Hey #PPCchat whats the deal with minimum send costs being mandatory in Google Ads even for single super cheap items? Google enforcement does not seem to be consistent

Reply on Twitter 2099117118390223074 Retweet on Twitter 2099117118390223074 0 Like on Twitter 2099117118390223074 0 Twitter 2099117118390223074
tweet_to_crazy avatar Anand Soni @tweet_to_crazy ·
11 Sep 2098458800676102196

𝗕𝗲𝗳𝗼𝗿𝗲 𝘆𝗼𝘂 𝘁𝗮𝗸𝗲 𝘁𝗵𝗮𝘁 𝗣𝗺𝗮𝘅 𝗥𝗢𝗔𝗦 𝗻𝘂𝗺𝗯𝗲𝗿 𝘀𝗲𝗿𝗶𝗼𝘂𝘀𝗹𝘆

ask yourself - have you excluded brand from it

#googleads #ppc #ppcchat #brandexclusion #pmax

Reply on Twitter 2098458800676102196 Retweet on Twitter 2098458800676102196 0 Like on Twitter 2098458800676102196 0 Twitter 2098458800676102196
Load More

Copyright © 2026 PPCChat | All Rights Reserved.