Skip to main content

A/B testing: real-world examples and what you actually gain

Author: Milos ZekovicReading time: 7 min

What A/B testing delivers in practice after years of running experiments: realistic limitations, traffic requirements, common mistakes, and the types of changes worth testing.

A/B testing: real-world examples and what you actually gain

What A/B testing is, and what it is not

A/B testing means showing one group of visitors version A and another group version B, then comparing a predefined outcome: a purchase, signup, form submission, demo request, or another conversion you can already track reliably.

I have been running and reviewing experiments on client websites since around 2019. The tools have changed since then, and the shutdown of Google Optimize left a gap. Today, teams use platforms such as VWO and Optimizely, server-side assignment, or their own custom solutions, while GA4 is typically used to analyze the results rather than display the variants.

The core principle has remained the same: one clear hypothesis, one primary metric, and enough traffic to reach a meaningful conclusion.

A/B testing is not:

  • a substitute for an unclear offer or broken checkout
  • a way to justify a redesign before addressing basic UX problems
  • a guarantee that every button-color change will increase conversions

It is a structured way to reduce guesswork on pages that already receive enough traffic and conversions for meaningful measurement.

What you actually gain

A clear answer to a specific question

Instead of debating whether someone prefers the blue or green button, you get an answer based on how visitors actually behave.

For example: under the conditions of this particular test, did variant B produce a higher form submission rate than variant A?

Alongside the primary metric, you can define additional measures to make sure the change does not create a problem elsewhere. These might include bounce rate, average order value, or the number of support requests.

Many tests do not produce a clear winner. That result is still useful because it can prevent you from launching a change that looked better in a meeting but did not improve real user behavior.

Priorities shaped by evidence

Teams that test regularly build a knowledge base of what worked, what did not, and under which conditions.

The outcome may depend on the audience segment, device type, season, traffic source, or overlap with marketing campaigns. Over time, priorities begin to shift from personal opinions toward documented results and lessons.

Limitations you should plan for

  • Traffic: reliable conclusions require enough conversions per variant. On smaller websites, a test may need to run for several weeks before it collects enough data. That is normal and does not mean the method has failed.
  • Privacy and tracking: cookie consent requirements and platform restrictions make precise tracking and segmentation more difficult than they were a few years ago. Server-side solutions and first-party analytics can help, but perfect tracking is unrealistic. Experiments should be designed with that limitation in mind.
  • Implementation cost: every variant must be built, tested, and connected to analytics correctly. Running ten micro-tests at the same time may sound efficient, but it usually complicates the analysis and adds noise to the data.

Patterns worth testing, without invented uplift numbers

These are types of experiments that often make sense once the fundamentals are solid: pages load quickly, forms work correctly, the mobile experience is usable, and the next step is clear.

I am not quoting potential percentage increases because the outcome depends on the traffic, offer, audience, and quality of the existing page.

E-commerce: CTA copy and shipping clarity

You might test the standard “Add to cart” treatment against a version that also makes the actual free-shipping threshold or expected delivery time more visible.

Any such promise must, of course, be accurate and consistent with the terms of purchase.

This type of test can reveal whether uncertainty about shipping, rather than the visual design, is preventing people from clicking. Sometimes clearer copy wins. Other times, the test simply reveals that the offer itself is not clear enough.

SaaS and B2B: a more specific opening headline

A headline that clearly describes an outcome, such as “Fewer missed deadlines on client projects,” can be compared with a broader category description such as “Project management software.”

This tests whether visitors understand who the product is for and which problem it solves as soon as the page loads.

Shorter is not automatically better, but a more specific headline is often easier to understand.

Lead generation: form length and field order

You can test removing fields that the sales team can complete later or moving the most sensitive question until after the visitor has entered an email address.

This helps you understand whether people are being discouraged by the length of the form or whether the offer itself is not relevant enough.

If form submissions increase but lead quality falls, your additional metrics should reveal it. More submissions do not necessarily produce a better business result.

Pricing and plan emphasis

The order of pricing plans, the option highlighted by default, and the way the main benefit is presented all influence how visitors make a decision, not just how the page looks.

Highlighting the middle plan only helps when that plan genuinely suits the majority of customers. Otherwise, it may create additional confusion or reduce trust.

What makes the results unreliable

When reviewing stalled or abandoned testing programs, I most often encounter the following problems:

  • Too many changes at once: if you change the headline, image, CTA, and content layout together, you will not know what actually affected the result.
  • Repeated peeking and early stopping: checking the results every day and ending the test as soon as one variant appears to be winning increases the risk of reaching the wrong conclusion.
  • The novelty effect: a major visual change can produce a short-lived lift simply because people notice it, but the effect often fades.
  • Uneven allocation or broken tracking: the result is unreliable if the variants do not receive comparable traffic or conversions are not measured consistently.
  • Testing before fixing the fundamentals: a slow LCP, broken mobile menu, or vague value proposition matters more than the color of a button.

Before launching a test, define the primary metric, minimum duration or required sample size, additional measures that protect the quality of the result, and the person responsible for deciding whether the variant should be launched or rejected.

Where A/B testing fits in the optimization process

On the websites I audit, experimentation works best as the final layer of optimization on pages that already have a stable, measurable conversion baseline:

  1. Clarify the offer and target audience on key landing pages
  2. Remove technical friction, such as slow loading, broken forms, and poor mobile usability
  3. Add credible trust signals, such as reviews, team information, and clear policies
  4. Only then test headlines, CTAs, form variants, and the placement of key elements

If the first three steps are weak, button-color tests usually distract from the real problem.

Tooling in brief

Since Google Optimize was discontinued, the right setup has depended on the budget, traffic volume, and engineering capacity of each team.

Platforms such as VWO, Optimizely, and Kameleoon, as well as custom server-side solutions, can be used to create and distribute variants. GA4, often combined with BigQuery, can support analysis and reporting.

The choice of tool matters less than how the testing process is managed: one clear hypothesis per test, consistent conversion definitions, and properly documented goals, changes, and results.

The same lesson has repeated itself since 2019: discipline matters more than the number of features in a dashboard.

When A/B testing makes sense

Run an A/B test when you have enough traffic and conversions, a hypothesis that can be confirmed or rejected, and a page where even a small improvement would justify the implementation time.

Skip it when the offer is unclear, tracking is unreliable, or the team is using a redesign to avoid addressing more important UX problems.

Used properly, A/B testing leads to fewer debates, clearer priorities, and occasional meaningful improvements. Tests that show no difference are equally useful because they can prevent you from launching a change that does not help users or the business.

Want to find out what could genuinely improve your website’s results?

Get in touch. We can identify the main conversion bottlenecks, define meaningful hypotheses, and create a reliable foundation for testing instead of making random changes.

Newsletter with ideas that matter

Subscribe to my newsletter. In the newsletter, I’ll share new insights, practical tips, and occasional case studies, everything that can help your business grow.

No spam, once a week or only when there’s something worth saying, and worth reading.