The A/B test begins before launch.

If the team first runs two versions and then decides which metric to look at, it’s tempting to choose the metric where the desired option looks better. Therefore, the hypothesis and the basic metric are fixed in advance.

Format of the hypothesis

“If we shorten the first screen on mobile KZ, then click → registration will grow because the CTA will be visible sooner.”

There's change, segment, metric and explanation.

One big variable.

If option B changes the title, shape, image, and order of the blocks at the same time, you’re testing the whole concept. This is acceptable for a large redesign test, but it does not help to understand which element is responsible for the result.

Why You Shouldn't Stop the First Leader Test

At the beginning of the sample, several random events are able to move the percentage strongly. Today, B can lead to +20%, tomorrow it will equal. The smaller the volume, the more careful the withdrawal.

What to watch besides the main metric

Basic metricGuardrail
click → registrationregistration → FTD
CTRpost-click CR
cost per FTDapproval rate

Guardrail is needed to ensure that the local victory does not worsen the next stage.

Segments after the test

First determine the overall result, then look at the larger segments. If B only wins on desktop and mobile traffic is 80%, a total rollout could be a mistake.

Document Losing Tests

Losing is also knowing. If the hypothesis does not work, write down the conditions. This saves the budget for repeating the same idea in two months.

When A/B is not possible

With a very small volume, it is better to use sequential observations and large changes, honestly acknowledging the limitations of the inference. It is not necessary to show statistical accuracy where there is no data.

Related material: How to Test a Landing.

Determine the minimum significant effect

Before the test, decide what difference will really change the business decision. If a 0.1 percentage point rise in CR doesn’t pay off the difficulty of supporting a new version, a victory of this magnitude is uninteresting even with statistical persuasiveness.

Don’t mix the test with traffic redistribution

If the source or GEO mix changes at the same time during A/B, the groups become less comparable. If possible, randomize options within the same stream.

Look at sustainability.

After the winner rollout, check to see if the full effect is maintained. Sometimes the experimental group was smaller and better quality, and after scaling the difference is reduced.

A negative result also closes the question.

If the test was correct and there is no significant improvement, this is an excuse to leave the current version and direct the resource to another hypothesis. You don’t need to “squeeze” an idea endlessly just because time has already been invested in it.

Fix technical failures separately

If during the test, one option for several hours gave an error, such a period can not simply be left in the final statistics. Mark the incident and decide in advance how to rule out corrupted data.

Don’t change the rules of victory along the way.

If the main metric was click → registration, you can not declare the winner after losing the time option on the page. Secondary indicators explain the result, but should not retroactively change the criterion of success.

Rollback should be simple.

Before the test, make sure the old version can be returned quickly. This is especially important for technical experiments on a form or redirect, where an error immediately affects traffic.