Skip to main content
Back to Curriculum
Module: Metrics, Growth & Experiments•Lesson 45•40 min read

A/B Testing & Experimentation

Lesson 45: A/B Testing & Experimentation

Three consecutive lessons have now ended by pointing forward to this one. Lesson 43's funnel Case Study found a plausible, evidence-based cause for a drop-off but stopped short of proving the proposed fix would actually work. Lesson 44's retention Case Study and Reflection Exercise both noted that a before-and-after cohort comparison cannot fully rule out confounding factors. In both cases, the missing piece is the same: a rigorous way to establish that a specific change causes a specific improvement, rather than merely correlating with one — precisely the correlation-versus-causation gap Lesson 41 first flagged and explicitly left unresolved until now.

This lesson closes that gap. A/B testing (more formally, controlled experimentation) is the discipline of randomly assigning users to different versions of a product experience and measuring the difference in outcomes between them — random assignment being the specific mechanism that allows a PM to conclude, with quantified confidence, that an observed difference was actually caused by the change rather than by some other factor that happened to vary alongside it. This lesson also covers the discipline's most common failure modes, since a poorly run experiment can produce a confident, precise-looking, and completely wrong answer — often more dangerous than no experiment at all, because it carries false authority.

Learning Objectives

  1. 1

    Explain why random assignment is the specific mechanism that allows a controlled experiment to establish causation, resolving the correlation-versus-causation gap from Lesson 41.

  2. 2

    Define statistical significance and explain, in plain language, what a p-value does and does not tell you.

  3. 3

    Explain sample size and statistical power, and why an underpowered experiment risks missing a genuine effect entirely.

  4. 4

    Diagnose "peeking" — stopping an experiment early based on an interim result — and explain why it inflates false-positive rates.

  5. 5

    Apply a pre-registration discipline (defining hypothesis, primary metric, guardrails, and sample size before launching) to design a trustworthy experiment.

This lesson assumes Lesson 41's full toolkit: precise metric definitions, Goodhart's Law, guardrail metrics, and especially the correlation-versus-causation caution this lesson directly resolves. It also assumes Lesson 43's and Lesson 44's open threads — both lessons identified a plausible fix for a real problem but explicitly deferred the question of rigorously validating that fix to this lesson.

Why Random Assignment Establishes Causation

Recall Lesson 41's caution: two metrics moving together doesn't establish that one causes the other, because a confounding variable might independently drive both. A controlled experiment solves this specific problem through random assignment: users are randomly split into a control group (seeing the existing experience) and one or more treatment groups (seeing the proposed change), with randomization ensuring that, on average, the two groups are statistically identical in every other respect — same mix of acquisition channels, same mix of device types, same mix of user tenure, same mix of literally everything else that might otherwise confound the comparison. Because random assignment neutralizes every other systematic difference between the groups, any statistically reliable difference in outcomes between them can be attributed to the one thing that was deliberately varied: the change being tested.

Process diagram showing flow: Eligible users → Random assignment → Control Group (existing Experience) → Treatment Group (proposed Change) → Measure outcome...

Yes

No

Eligible users

Random assignment

Control Group (existing Experience)

Treatment Group (proposed Change)

Measure outcome

Statistically Reliable
Difference Between Groups?

Change Likely Caused
the Observed Difference

No Reliable Evidence
the Change Had an Effect

Statistical Significance and P-Values, in Plain Language

A p-value answers a specific, narrow question: if the change genuinely had no real effect at all, how likely would it be to observe a difference this large (or larger) between groups purely by random chance? A small p-value (conventionally, below 0.05, though the right threshold depends on context) suggests the observed difference would be unlikely to arise from chance alone if there were truly no effect, giving some confidence that a real effect exists. Critically, a p-value does not tell you the probability that the change actually works, nor does it tell you how large or practically meaningful the effect is — a large sample size can produce a statistically significant result for an effect too small to matter practically, and a small sample size can fail to reach significance even for an effect that's genuinely large and meaningful, simply because there wasn't enough data to detect it reliably.

Sample Size and Statistical Power

Statistical power is the probability that an experiment will correctly detect a real effect, if one genuinely exists, at a given sample size. An underpowered experiment — one run with too few users relative to the size of the effect being tested for — risks concluding "no significant difference found" not because the change had no effect, but simply because the experiment never had enough data to reliably detect an effect of that size, even if a real one existed. Before running an experiment, a PM should estimate the minimum detectable effect (the smallest change worth caring about, business-wise) and ensure the planned sample size is large enough to reliably detect an effect of that size — running an experiment without this calculation risks either wasting resources on a hopelessly underpowered test, or running it far longer than necessary for an effect large enough to be obvious much sooner.

The Peeking Problem

A specific and extremely common mistake deserves detailed treatment: peeking — checking an experiment's results before it reaches its planned sample size or duration, and stopping it early the moment the result happens to look statistically significant. This inflates the true false-positive rate far beyond the nominal significance threshold, because random noise naturally causes an experiment's measured difference to fluctuate above and below the true effect throughout its run — checking repeatedly and stopping at the first moment the fluctuation happens to cross a significance threshold is systematically biased toward catching noise, not signal, since a sufficiently long-running experiment on a null effect will, by chance, cross a nominal 0.05 significance threshold at some point during its run far more than 5% of the time if checked and potentially stopped repeatedly.

Process diagram showing flow: Experiment starts → Day 1: Check Result— Not Significant Yet → Day 2: Check Result— Not Significant Yet → Day 3: Check Result— Appears Significant! → Stop and Ship Based on Day 3 Result...

Risk: This May Be a Random
Fluctuation, Not a Real Effect

Experiment starts

Day 1: Check Result
— Not Significant Yet

Day 2: Check Result
— Not Significant Yet

Day 3: Check Result
— Appears Significant!

Stop and Ship Based on Day 3 Result

False positive

The correct discipline is to determine the required sample size and duration before launching the experiment (per the previous section), and either wait until that pre-determined point to analyze results, or use a statistical method specifically designed to allow valid early stopping (sequential testing methods), rather than informally checking and stopping at the first appealing-looking result.

Pre-Registration: Deciding Before You Look

The most reliable defense against both peeking and other, subtler forms of unintentional bias (sometimes called p-hacking when done more deliberately) is pre-registration: writing down, before the experiment launches, the specific hypothesis being tested, the single primary metric that will determine success, the guardrail metrics (Lesson 41) being monitored to catch unintended harm, the planned sample size and duration, and the specific threshold that will count as a meaningful result. Committing to these decisions in advance prevents the natural, often unconscious temptation to retroactively decide "well, this secondary metric moved significantly, so let's call that our success metric instead" after seeing results that didn't confirm the original hypothesis — a pattern that, applied loosely enough across many possible metrics, will eventually find something that looks significant purely by chance, without that finding representing a genuine effect.

Common Mistakes to Avoid

✕

Treating a statistically significant result as proof the change definitely works and is worth shipping

As covered in Theory, statistical significance indicates the observed difference is unlikely to be pure chance — it says nothing about whether the effect size is practically meaningful, and a significant-but-tiny effect may not be worth the cost of shipping and maintaining the change.

✕

Running an experiment without first calculating the required sample size

As covered in Theory, this risks either wasting effort on a severely underpowered test that will very likely fail to detect a real effect, or leaving an experiment running far longer than necessary once a large, obvious effect could have been detected sooner.

✕

Peeking at results and stopping the experiment early upon seeing an appealing result

As covered in Theory, this specific behavior systematically inflates false positives, and is one of the single most common ways a well-intentioned experimentation program produces unreliable, non-reproducible results.

✕

Checking many secondary metrics after the fact and treating whichever one moved significantly as the "real" result

This is a subtler cousin of peeking — testing enough metrics, by chance alone, some will appear significant even with no genuine underlying effect, which is precisely why pre-registering a single primary metric in advance is essential.

✕

Ignoring guardrail metrics because the primary metric improved

An experiment that improves its primary metric while quietly damaging a guardrail metric (echoing Lesson 41's Goodhart's Law caution, and Lesson 45's own version of the same principle applied to a single experiment rather than an ongoing target) should not be shipped without understanding and addressing that trade-off explicitly.

Mental Model

The Peeking Trap

This lesson's core takeaway tool visualizes why informal, repeated checking of an in-progress experiment is fundamentally different from — and far riskier than — a single, planned analysis at a pre-determined endpoint:

Use the Peeking Trap as a standing discipline whenever tempted to check an in-progress experiment's results before its planned endpoint: remind yourself that an appealing-looking interim result is not evidence of a real effect — it's exactly the kind of noise a properly powered, pre-registered experiment is specifically designed to filter out, and giving in to the temptation to stop early defeats that purpose entirely.

Quick Reflection Checkpoint

Key Takeaway: How will you apply "The Peeking Trap" when evaluating trade-offs in your product decisions?

Ready to test your product judgment?

Take the interactive practice quiz for Lesson 45 and build your skill radar dashboard.