A/B testing landing pages is the most reliable way to learn whether a change caused more enquiries, rather than arriving alongside a good month or a new campaign. It is also widely misused. Run on too little traffic, a test doesn’t just fail to find an answer; it can hand you a confident answer that is wrong.
So the first question isn’t what to test, but whether your page can support a test at all, and what to do if it can’t.
The short answer
A/B testing, also called split testing, sends visitors at random to two versions of a page and compares how many in each group convert. Because both groups meet the same season, ads and conditions, a difference too large to be chance can be credited to the change.
- It works when the page collects enough conversions to detect the lift you care about: at the usual 95% significance and 80% power, roughly 400 per version for a 20% relative lift and 1,600 for a 10% lift.
- It fails when tests stop early, use the wrong metric or run on a few dozen conversions a month.
- Without the traffic, fix obvious problems, research with real people, test bold changes rather than tweaks, and try messages in ads first.
How A/B testing works, in plain terms
A testing tool assigns each visitor to version A (the control, your current page) or version B (the variant), shows returning visitors the same version, and counts conversions in each group. Two relatives are worth knowing:
- Split URL testing sends traffic to two separate pages at different addresses, which suits a redesigned page.
- Multivariate testing tests every combination of several changes. Two headlines, two images and two buttons make eight versions, each needing its own share of traffic, so it only suits very busy pages.
The four numbers behind every test
Sample size calculators ask for four inputs:
- Baseline conversion rate: how the page converts for the traffic you’ll test, from GA4 over the last four to eight weeks.
- Minimum detectable effect (MDE): the smallest improvement worth detecting, usually as a relative lift. Going from 3% to 3.6% is a 20% lift.
- Significance level: conventionally 5%, the source of “95% confidence”. It limits how often you’ll crown a winner when there is no real difference.
- Statistical power: conventionally 80%, the chance of detecting a real improvement as large as your MDE.
What statistical significance means. A result significant at 95% says that if both versions performed identically, a gap at least this large would appear less than 5% of the time. It isn’t a 95% chance that B is better, and on its own it doesn’t tell you how big the improvement is.
Check your A/B test sample size before you build anything
Run the numbers before designing a variant. At typical landing page conversion rates, the conversions you need per version depend mostly on the lift you want to detect, not on the conversion rate itself. Here is a standard calculation at 95% significance and 80% power, with visitors shown for a page converting at 3%:
| Relative lift to detect | Conversions per version (approx.) | Visitors per version at 3% |
|---|---|---|
| 10% (3.0% → 3.3%) | 1,600 | 53,000 |
| 20% (3.0% → 3.6%) | 420 | 14,000 |
| 30% (3.0% → 3.9%) | 190 | 6,500 |
| 50% (3.0% → 4.5%) | 75 | 2,500 |
For the total, double the per-version figure in an A/B test, and add another version’s worth for every extra variant. Calculators use slightly different formulas, so expect similar rather than identical answers.
A worked example
A landing page receives 4,000 paid visitors a month and converts at 3%: 120 enquiries.
- Detecting a 20% lift needs about 28,000 visitors in total: about seven months.
- Detecting a 50% lift needs about 5,000: five to six weeks.
So this page can test big differences, such as a new offer, but not a new button colour. Long tests are also fragile: the longer one runs, the more seasons, budget shifts and visitors clearing cookies (and being reassigned) disturb it.
Even when the sample arrives quickly, run for at least two full weeks, because weekday and weekend visitors often behave differently.
Why stopping early produces false winners
A common route to a wrong answer is checking results daily and stopping the moment the tool shows a winner.
The peeking problem
Early on, counts are small and the gap between versions swings widely. A standard significance test assumes you look once, at the planned sample size. Stop at the first significant reading and you give random noise many chances to cross the line, so the real false-positive rate ends up far above 5%.
The fix is dull: set the sample size and duration before launch, and don’t act until you reach them. Some tools use sequential statistics designed for monitoring as data arrives; if yours does, follow its stopping rules exactly.
Other ways tests mislead
- Sample ratio mismatch. You set 50/50 and got 55/45 on thousands of visitors. Something, perhaps a redirect, is losing visitors. Don’t trust the result.
- The wrong metric. Button clicks rise while completed enquiries don’t. Choose one primary metric in advance and make sure it is tracked properly in GA4.
- Too many comparisons. Check ten metrics or slice by every device and country, and something is likely to look significant by chance.
- The winner’s curse. Winners from underpowered tests tend to overstate the effect. Expect less once the change is live.
- Changes mid-test. A new ad, budget change or page edit alters who arrives and what they see, muddying the result.
Before calling a winner, ask sales whether the extra enquiries are genuine prospects.
What to A/B test on a landing page, and what to skip
Most pages can only detect large effects, so test what changes a visitor’s decision. Test one idea rather than one element: a variant can change a headline, image and subheading together if they express a single hypothesis.
| Worth testing | Why it can move conversions | Example |
|---|---|---|
| The offer | Changes what visitors get for acting | “Book a consultation” against “Get a fixed-price estimate” |
| Headline and value proposition | Decides whether visitors think the page is for them | Outcome-led against feature-led |
| Proof | Answers “can I trust this?” at the moment of decision | Project results beside the form against testimonials further down |
| Form length | Trades lead volume against lead quality | Three fields against seven, judged on qualified leads |
| Page length and order | Matches depth of explanation to the decision | A short page against one that answers objections |
Rarely worth testing at typical volumes: button colours, fonts, small image swaps and minor wording. Very large sites can detect lifts of a fraction of a percent; a page with 100 enquiries a month never will, so make the sensible choice and save your traffic.
For ideas, diagnose first. Why your landing page isn’t converting works through a page in the order visitors meet it, and the conversion rate optimisation guide turns findings into prioritised hypotheses.
A/B testing landing pages without harming SEO or speed
If the page appears in Google, follow Google Search Central’s guidance on minimising A/B testing impact in Search:
- Don’t cloak. Never show Googlebot one version and people another. Google treats cloaking as a spam policy violation.
- Use
rel="canonical"on variant URLs, pointing to the original page, rather thannoindex. - Use 302 (temporary) redirects, not 301s, in split URL tests, so Google keeps the original URL.
- End tests promptly, then remove variant URLs, redirects and test scripts.
Watch the performance cost
Client-side testing tools change the page in the browser after it starts loading. To stop visitors glimpsing the original first (known as flicker), many hide the page until the test script has run. That can delay Largest Contentful Paint, which Google rates as good at 2.5 seconds or less. Run both versions through PageSpeed Insights before launch, using your tool’s preview link for each: if one is heavier, you’re testing speed as well as the idea.
Testing tools now Google Optimize has gone
Google Optimize closed in September 2023, and GA4 doesn’t run experiments itself. At the time of writing (June 2026), the options fall into four groups:
- Client-side platforms with visual editors: quick for marketers, but prone to flicker.
- Server-side and feature-flag platforms: no flicker, but they need developer time.
- Landing page builders and CMS plugins that split traffic between page versions.
- Ad platform experiments, such as those in Google Ads and Meta.
Whichever you use, judge versions on the enquiries your website KPIs already track.
A/B testing on low-traffic websites: what to do instead
Many service and B2B landing pages will never have the traffic for tests like these. That isn’t a reason to guess, but to use methods that suit the volume.
Fix what doesn’t need proving
A form that fails on some phones, an eight-second load, a headline that doesn’t say what you sell. None of these needs an experiment. Fix them and note the date.
Research with real people
Qualitative research shows why people hesitate, which a test never does. Jakob Nielsen’s widely cited Nielsen Norman Group article argues that testing with around five users finds most usability problems, and that several small rounds of testing beat one large study.
- Watch five people from your audience try to act on the page, on a phone, thinking aloud.
- Ask sales what prospects ask before buying, and check the page answers it.
- Ask new enquirers: “What nearly stopped you getting in touch?”
Run five-second tests
Show someone the page for five seconds, hide it, then ask what it offers, who it’s for and what they would do next. If they can’t say, the first screen isn’t working. It needs no traffic at all.
Compare before and after, with caveats
A change made without a test can still teach you something:
Even then, you can’t fully separate your change from everything else that moved, so read a large, sustained shift as a good sign and a small one as noise.
Test the message in ads first
Run ads making genuinely different promises, such as speed, price certainty or expertise, to one audience, then lead the landing page with the winner. Delivery tends to favour whichever ad does well early, so use the platform’s experiment feature for a fair split. Clicks aren’t enquiries, so check which ad’s visitors went on to enquire.
Make bigger swings, or pool traffic
A bold alternative, such as a different offer, can be tested on traffic that would never detect a tweak. If several campaign pages share a template, testing a change across all of them pools their traffic.
Frequently asked questions
How long should an A/B test run?
Until it reaches the sample size calculated before launch, over at least two whole weeks. If that means more than a couple of months, test a bolder change instead.
Does A/B testing hurt SEO?
Not when it follows Google’s guidance: no cloaking, canonical tags on variant URLs, temporary redirects and ending tests promptly. The bigger practical risk is speed, because testing scripts can slow the page.
Can I trust the “chance to win” figure in my testing tool?
Treat it as one input, not a verdict. Each tool calculates it under its own assumptions, and it is unreliable on small samples or if you stop the moment it looks good.
What to do next
A/B testing is a precise instrument with a big appetite for traffic. Use it where the numbers allow, on decisions that matter. Everywhere else, fix what’s broken, listen to real visitors and track carefully enough that a before-and-after comparison means something.
Both start with a page that is fast and focused on one action. Our landing pages are built that way. For an outside view of a page you already run, a free website audit covers speed, mobile experience and how easily visitors can get in touch.