A feature on letting the data decide
Two versions enter. The data decides.
A/B testing is the controlled experiment at the heart of conversion work — the method that lets two versions of a page compete on real, live traffic so the better one is proven, not merely argued for. It is the closest thing marketing has to a laboratory: split the audience, change one thing, measure the result, and let the evidence settle what no meeting ever could. Done with rigour it is the most honest tool you have. Done carelessly — stopped early, run too small — it will cheerfully tell you comforting lies.
Two page variants labelled A and B side by side, with a results chart between them showing a confidence interval and a clearly proven winner. The controlled experiment that settles the question. Square aspect ratio, dark studio, focused light, solar-yellow graph glow, a sense of rigour and proof.
Two versions, real traffic, one proven winner — the experiment that turns a confident opinion into an actual fact.
Every confident claim about what will work on a website is, until tested, a guess wearing a suit. The headline that "obviously" reads better, the button colour someone feels strongly about, the longer page versus the shorter one — these are arguments nobody in the room can actually win, because none of them is the customer. A/B testing ends the argument by running it as an experiment. Instead of debating which version is better, you show both to real visitors, at random, and let their behaviour declare the winner. The opinion becomes a hypothesis, and the hypothesis gets tested.
The mechanics are simple to state and surprisingly easy to get wrong. You take a single, well-formed hypothesis — change this one thing, and this specific number will improve, for this reason — and you split your incoming traffic at random between the current version and the new one. Everything else is held constant. You measure one agreed success metric, and you keep running until enough visitors have seen each version for the result to be trustworthy. The discipline is in the constraints: one change, one metric, a fair split, and the patience to wait for a real answer.
Where A/B testing goes wrong is almost always statistical, and almost always in the same way: calling a winner too soon. A test that looks like it's winning after a day is, very often, just noise that hasn't averaged out yet — and a team that stops the moment the numbers look good will "prove" a long string of improvements that quietly evaporate on rollout. Peeking and stopping early, samples too small to mean anything, judging on the wrong metric: these are the ways a test lies. Getting them right is the entire difference between testing and self-deception.
In this feature
Six things serious A/B testing insists on.
A real hypothesis
A good test starts with a clear statement: change this, and this metric will move, for this reason. The hypothesis is what makes a result meaningful — so we never test a random tweak, only a specific, reasoned bet we can actually learn from.
A fair split
Visitors are divided at random between versions, with everything else held constant, so the only difference is the one being tested. A clean split is what lets us attribute the result to the change rather than to chance or some hidden variable.
One success metric
A test needs one agreed measure of success, decided before it starts. We define the real goal — orders, sign-ups, revenue — not a flattering proxy, because a "win" on clicks that doesn't move the metric that matters is no win at all.
Real significance
We run each test to a pre-agreed sample size and proper statistical significance before reading it. That patience is what separates a genuine, repeatable lift from a flattering blip — and it is the single most common thing amateur testing gets wrong.
No peeking
Stopping a test the moment it looks like winning is how you "prove" gains that vanish on rollout. We set the stopping rules in advance and hold to them, so the result is honest — not a number caught at its most flattering moment.
Win, lose or learn
A losing test is not a failure — it is a saved mistake and a lesson about your customers. We report results honestly, banking the wins and learning from the rest, because a programme that only ever "wins" isn't testing; it's fooling itself.
Two versions.
One truth.
The reason A/B testing matters so much is that humans, including expert ones, are bad at predicting what other humans will do on a page. The version everyone in the room loves loses surprisingly often; the change someone dismissed as trivial sometimes moves the number more than the grand redesign. This isn't a knock on anyone's judgement — it is simply that taste and conversion are different things, and only the audience can tell you which version actually performs. Testing is how you stop paying for the gap between what you think will work and what does.
A clean experiment depends on changing one thing at a time. If you test a new headline, a new image and a new button all at once and the version wins, you have learned that something worked but not what — and you cannot reliably reuse the lesson. Isolating the variable is what turns a result into knowledge: not just "this page did better" but "this specific change caused it." That discipline is slower than redesigning everything and hoping, and it is the difference between a test that teaches you about your customers and one that merely produces a number.
The deepest discipline, though, is statistical patience, because this is where almost everyone slips. Early in a test the numbers swing wildly, and a version can look like a decisive winner purely by chance before enough visitors have evened things out. Stop there and you bank a phantom. The guard against it is unglamorous: decide the required sample size and significance level before you start, refuse to peek-and-stop, and accept that a trustworthy answer takes as long as it takes. A test you ended early isn't a fast result — it's an unreliable one.
The serious version of A/B testing treats every experiment as a way to learn, not just to win. A losing test is not wasted — it has saved you from shipping a change that would have cost you, and taught you something real about what your customers respond to. So we form a clear hypothesis, isolate the variable, run to genuine significance, and report the outcome honestly whichever way it falls. Over time that builds something more valuable than any single lift: a tested, evidence-based understanding of your audience that guesswork can never match.
A clear A/B test results screen — variant A versus variant B, a confidence interval, a winner marked at high significance. The moment an experiment delivers a trustworthy verdict. Wide cinematic 21:9 crop, dark elegant studio backdrop, solar-yellow accent, a sense of rigour and proof.
One change, one metric, a fair split, real significance — the rigour that turns a confident opinion into a proven fact.
Five questions we ask before calling a winner.
The version everyone in the room loves loses surprisingly often. Only the audience can tell you what actually works — so the whole job is to ask them properly, and listen to the answer.
A/B testing is the engine room of conversion work. It is the method beneath CRO's strategy, the way UX hunches about a better flow are proven, and the tool that turns landing-page and web-design choices from matters of taste into matters of fact. Because our testers sit beside the designers and developers, a proven winner doesn't languish in a report — it ships, and the next hypothesis is already queued. Strategy decides what to test; A/B testing decides what is true.
We run experiments with real statistical rigour — one variable, one metric, a fair split, genuine significance — and we report every outcome honestly, because a test you cannot trust is worse than no test at all. The brief is a verdict you can bank: a proven winner, or a saved mistake and a lesson learned. That discipline is exactly why A/B testing is the backbone of Conversion Engineering: it is the difference between a site shaped by evidence and one shaped by the most confident voice in the room.
A screen showing two product-page variants in a head-to-head test, the surprising winner highlighted against the one everyone expected. The moment a test overturns a confident assumption. Contemporary, shallow depth of field, dark desk, solar-yellow screen glow.
The version nobody backed won — the test settled it, and saved a confident redesign from going live.
Representative scenario · not a named client engagement
An e-commerce team was about to ship a redesign everyone loved — until a controlled test revealed it would have cost them.
The team had invested heavily in a bold new product-page design, and internal enthusiasm was high — it looked more modern, more on-brand, plainly "better." The plan was to roll it out across the store immediately. The only dissenting voice was a quiet one: how do we actually know it converts better? Rather than rely on the confidence in the room, they agreed to settle it the honest way, with a controlled experiment on live traffic before committing to the change.
We set up a clean A/B test: the existing page against the new design, traffic split at random, one success metric — completed purchases — agreed in advance, and a sample size and significance level fixed before we began. Then we did the hardest part: we waited, resisting the temptation to call it early when the numbers swung. We also isolated a separate, much smaller change to the page for its own test, to see what was really moving the needle.
The admired redesign converted slightly worse; a small, unglamorous tweak — tested separately — won clearly instead.
The result was humbling and hugely valuable. Rolling out the beloved redesign would have quietly lowered conversion across the whole store — a confident, expensive mistake, avoided. Meanwhile the small change nobody had championed became a real, banked win. The test didn't just find a winner; it overturned the conviction of the entire room with evidence — which is precisely why no significant change now ships there without one. Opinion proposes; the experiment decides.
We were completely sure our new design was better — and a Revolutionize A/B test proved it would have cost us money. They saved us from shipping a confident mistake and found the change that actually won instead. Nothing significant goes live here now without a test; opinions just start the conversation.
What serious A/B testing actually involves.
A/B testing is an investment in making decisions on evidence instead of opinion, and it works best as an ongoing capability rather than a one-off. Engagements range from designing and running a single high-stakes experiment to operating a continuous testing programme across your site, and the scope rises with the number of tests, the complexity of the experiment design, and the tooling and analytics involved.
The single biggest factor is traffic volume, because a test cannot reach significance without enough visitors — and that determines both which tests are worth running and how long each takes. A high-traffic site can run many experiments quickly; a lower-traffic one must pick its battles and test only changes large enough to detect. We are honest about this up front, because running underpowered tests that can never reach significance is worse than not testing at all.
Every engagement includes the full discipline: a clear hypothesis for each test, clean experiment design with one isolated variable, a single success metric agreed in advance, a proper sample-size and significance plan, honest analysis with no peeking, and clear reporting of every outcome — with winners built and shipped in-house by the same team.
We scope every engagement to your traffic and goals rather than to a price list — which is why we don't publish rate cards. Every one begins with a free 30-minute scoping conversation, and we will tell you honestly whether your traffic can support meaningful testing yet, or whether your effort is better spent elsewhere first. We would rather decline a test we can't make significant than charge you to chase noise.
When you're ready
Stop arguing. Run the test.
Tell us about the change your team can't agree on, or the assumption baked into your site that nobody has ever checked. We'll respond within 24 hours with an honest read on whether it can be tested properly on your traffic — and how we'd design the experiment to settle it for good.
Begin the conversation →