FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Web & Mobile ¡ 5 minute read

A/B Testing Implementation: Getting Results You Can Trust

An A/B test is only as good as its implementation. Stable assignment, a sample size decided in advance, one primary metric with guardrails, and a fixed stopping point are what separate a result you can act on from a number that looks precise and means nothing.

By FISTA Solutions¡ AI-Native Engineering Team¡
A/B Testing Implementation: Getting Results You Can Trust article cover

Most A/B test implementations produce numbers that look precise and mean very little, usually because of how assignment, stopping, or metric choice was handled. This guide covers getting results you can act on, drawing on FISTA Solutions' web and mobile work.

What does a valid test require?

Five decisions, all made before the test starts.

DecisionWhy it matters before the start
Primary metricPrevents choosing the winner after the fact
Minimum effect sizeDetermines whether the test is worth running
Sample sizeFixes the stopping point
DurationCovers weekly cycles and novelty effects
GuardrailsDefines what would veto a win
Assignment unitUser, session, or account

How should assignment work?

Hash a stable identifier into a bucket, so a given user always sees the same variant.

Per-request assignment produces users who see the feature intermittently. That is confusing to them and fatal to the analysis, because the treatment and control groups are no longer distinct populations.

Choose the assignment unit to match the effect. If users collaborate inside an account, assign by account, or the variants leak into each other. See feature flags guide.

Why is peeking so damaging?

Because checking repeatedly and stopping when the result crosses a threshold inflates the false positive rate well beyond what the threshold implies.

A test designed for a five percent error rate, checked daily and stopped on the first significant reading, produces false positives far more often. The threshold assumes one look at a predetermined sample size.

Either fix the sample size and look once, or use a sequential method designed for repeated looks. Do not use fixed-sample statistics and look every morning.

How do you choose a sample size?

From the effect size you would actually act on, not from what you hope to see.

Work backwards: if a two percent relative improvement would change your decision, calculate the sample needed to detect two percent. If that exceeds your traffic for the next quarter, the test is not worth running.

This calculation is where most experimentation programmes should stop and frequently do not. An underpowered test does not produce a weak signal; it produces noise that looks like a signal.

What should you measure?

One primary metric that decides the test, and a small set of guardrails that can veto a win.

Tracking many metrics and reporting whichever moved is how teams convince themselves of effects that are not there. With twenty metrics, something will look significant by chance.

Guardrails matter because a change can improve one number by damaging another. A conversion gain that also raises support contacts and page load time is usually a net loss. See observability for web apps.

How long should a test run?

At least one full weekly cycle, and long enough to pass any novelty effect.

Behaviour differs across days of the week. A test that runs Tuesday to Thursday measures Tuesday-to-Thursday users, which is a different population from your actual one.

Novelty matters for interface changes: users react to difference before they react to quality. A week or two of exposure separates the two.

How do you avoid instrumentation bugs?

By running an A/A test before you trust the system.

Assign users to two identical variants and confirm the metrics come out equal. If they do not, the problem is in assignment, logging, or analysis, and every subsequent result is suspect.

Also check sample ratio: if you assigned fifty-fifty and observe fifty-three to forty-seven, something is wrong with assignment or with data loss, and the result should not be used.

What are the common mistakes?

Per-request assignment. Stopping when the numbers look good. No minimum effect size. Many metrics and no primary. Tests too short to cover a week. And no A/A validation of the pipeline.

How do you test it?

Run an A/A test first. Check sample ratio on every experiment. Verify that the variant a user sees matches what the logs record, because a mismatch invalidates everything downstream.

What does it cost to operate?

A managed experimentation platform carries a subscription; a basic implementation is straightforward and the analysis is the hard part either way.

The real cost is opportunity: an underpowered test consumes weeks of traffic and produces no decision.

What should you measure?

Tests run versus tests that produced a decision, proportion stopped early, sample ratio mismatches detected, and how often a shipped winner's effect persisted when checked later.

How does this apply to AI features?

Directly, with one addition: quality metrics are usually offline while business metrics are online. An evaluation suite tells you whether outputs improved; a live test tells you whether users behaved differently, and the two frequently disagree.

Run both. Ship on the online result, but require the offline evaluation not to regress. See how to build an agent evaluation harness.

When is this the wrong approach?

With a few hundred visitors a week, no realistic test detects the effects you care about. Qualitative research and judgement are better tools at that scale, and pretending otherwise produces false confidence.

What should you do first?

Calculate the sample size needed to detect the smallest effect you would act on. If your traffic cannot reach it in a reasonable period, stop building the testing infrastructure.

How FISTA Solutions helps

FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: sample size and stopping point fixed before a test starts, an A/A run to validate the pipeline before any result is trusted, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To scope this work, message FISTA on WhatsApp, or read feature flags guide.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How should users be assigned to variants?

By hashing a stable identifier so the same person always lands in the same bucket. Per-request assignment means users flip between variants, which confuses them and makes the measurement meaningless.

02Why decide sample size in advance?

Because checking results repeatedly and stopping when they look significant produces false positives at a much higher rate than the stated threshold. Fixing the sample size and duration in advance is what makes the statistics valid.

03How many metrics should a test have?

One primary metric that decides the outcome, plus guardrails that can veto. Testing many metrics and reporting whichever moved is how teams convince themselves of effects that are not there.

04What are guardrail metrics?

Measures that must not degrade even if the primary metric improves — error rates, page speed, support contacts, unsubscribes. A conversion gain paid for with a latency regression is usually not a gain.

05When should you not A/B test?

When traffic is too low to detect the effect size you care about, when the change is obviously correct, or when the decision is strategic rather than incremental. Underpowered tests produce noise that looks like evidence.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project