Lessons · Lesson 2 of 3
Peeking, and the segment you went looking for
Measure what repeated looks do to a 5% false-positive rate, see how badly a test stopped early overstates itself, and treat a winning segment as the hypothesis it is.
Lesson 2 of 3 · 36 min
Two ways to find something that is not there
Two ordinary habits can turn a result of nothing into a result. Both feel like diligence rather than cheating.
The first is stopping an experiment on the morning it looks good. The second is accepting that it found nothing overall, then hunting group by group until you find somewhere it worked. This lesson measures what each habit does to your chances of being fooled.
Lesson 1's test found nothing, over a month, on 1,223 orders. Two things happened around it. Each would have produced a "finding" out of that same nothing, and both are so ordinary that most people do them without noticing.
The first was Thursday. Somebody looked at a running test, liked what they saw, and stopped it. The second happened the week after it ended. The result came back flat, so a colleague went looking for the segment where it had worked.
Neither is dishonesty. Both are ways of taking more chances than the arithmetic was priced for.
Peeking: why the fifth look is not free
A significance test is a bet with a stated house edge. Set the threshold at 1.96 standard errors and you have agreed to this: when the change does nothing at all, you will be fooled 5% of the time. That 5% is the price of a single look at a fixed sample.
A running test does not sit still. The gap between the two arms wanders as orders arrive. It goes up on a Tuesday and back down on a Wednesday. Meanwhile the standard error shrinks as the sample grows. Every time you look, you give that wandering path another chance to be over the line at the moment you happen to be watching. Stop the first time it is, and you have not run a test at 5%. You have run something else, and nobody in the room knows what.
The tempting way to price this is to multiply. Ten looks, each 5%, so one minus 0.95 to the tenth, which is 40.1%. That is wrong, and it is wrong in the direction that flatters you. Ten looks at a growing dataset are not ten independent tests. The day-14 numbers CONTAIN the day-7 numbers, so a path that was high on day 7 is more likely to still be high on day 14. The looks are heavily correlated, and correlated chances overlap.
So the number has to be measured rather than derived. Here is the measurement, and you can reproduce it.
| Looks during the test | Look every | Reported significant at some look | Significant at the final look only |
|---|---|---|---|
| 1, at the end | 25,200 sessions an arm | 5.01% | 5.01% |
| 4, weekly | 6,300 sessions an arm | 12.63% | 4.95% |
| 10 | 2,520 sessions an arm | 19.28% | 5.02% |
| 28, daily | 900 sessions an arm | 27.56% | 5.01% |
Read the last column first. It is the control on the experiment. However many times you LOOK, the final reading is still wrong 5% of the time, exactly as advertised. Looking does not corrupt the data. It corrupts the decision rule, and only if you act on what you see.
Now read the first row as the check that the simulation itself is sound. One look returns 5.01%, which is the 5% everyone already knew. A simulation that cannot reproduce the answer you already have is a bug, not a finding.
A team that watches a daily dashboard and ships whatever crosses the line is running at 27.56%, not 5%. More than one such "win" in four is nothing at all.
The stopped-early winner overstates itself, always
There is a second cost. It damages next year's plan rather than this week's decision.
If you stop the moment the line is crossed, you stop on a high. The estimate you carry out of the room is not a fair measurement of the effect. It is a measurement conditional on having been extreme enough to cross a threshold, which is a different and larger thing. The same simulation records what those stopped tests reported.
| Looks during the test | Median relative lift reported by the tests that stopped as winners |
|---|---|
| 1, at the end | +13.6% |
| 4, weekly | +20.5% |
| 10 | +25.5% |
| 28, daily | +33.3% |
Every one of those lifts is invented in the strictest sense. The two arms were identical by construction. A team that peeks daily and stops on its first winner will typically walk out claiming about a third more sales from a change that does nothing.
Now look back at lesson 1. Thursday's reading was +18.6%. That sits comfortably in the middle of what pure noise produces when you look every day. It is the most useful thing anyone could have said in that meeting.
Notice the first row too. Even a clean, single-look, properly-powered test overstates the effect when it wins. A result only crosses the line if it is at least as large as the detectable effect, so the winners are drawn from the top of the distribution. This is why a change that tested at +18% in a well-run experiment so often delivers less than that when it is rolled out. And it is why the rollout is not evidence that somebody sabotaged the launch.
The segment you went looking for
The test ended flat. The following Tuesday a colleague sliced the result by device and channel, twelve segments in all, to find out where the panel HAD worked. Here is the whole slice. The sessions and orders add back to the totals from lesson 1: 25,214 and 601 in the control, 25,180 and 622 in the variant.
| Segment | Control sessions / orders | Rate | Variant sessions / orders | Rate | Relative lift | p |
|---|---|---|---|---|---|---|
| Desktop, direct | 2,410 / 76 | 3.15% | 2,402 / 71 | 2.96% | -6.3% | 0.690 |
| Desktop, organic search | 3,180 / 92 | 2.89% | 3,166 / 95 | 3.00% | +3.7% | 0.800 |
| Desktop, paid search | 1,240 / 33 | 2.66% | 1,252 / 31 | 2.48% | -7.0% | 0.770 |
| Desktop, email | 1,020 / 36 | 3.53% | 1,014 / 40 | 3.94% | +11.8% | 0.621 |
| Phone, direct | 3,640 / 74 | 2.03% | 3,651 / 78 | 2.14% | +5.1% | 0.757 |
| Phone, organic search | 6,050 / 118 | 1.95% | 6,039 / 121 | 2.00% | +2.7% | 0.834 |
| Phone, paid social | 1,140 / 23 | 2.02% | 1,128 / 39 | 3.46% | +71.4% | 0.035 |
| Phone, email | 1,330 / 30 | 2.26% | 1,322 / 27 | 2.04% | -9.5% | 0.705 |
| App, direct | 1,860 / 59 | 3.17% | 1,848 / 62 | 3.35% | +5.8% | 0.754 |
| App, email | 640 / 24 | 3.75% | 646 / 22 | 3.41% | -9.2% | 0.739 |
| Tablet, all sources | 1,510 / 22 | 1.46% | 1,502 / 24 | 1.60% | +9.7% | 0.753 |
| Other and unattributed | 1,194 / 14 | 1.17% | 1,210 / 12 | 0.99% | -15.4% | 0.668 |
One row is under 0.05. The story writes itself: the fit panel works on phone, paid social. That is a cold audience arriving from an advertisement, with no idea what an Ellinghay last is. It is exactly the shopper a fit panel should help. It is a good story. It even has a mechanism, which is what makes it dangerous.
What that row is worth
Twelve segments, each given a 5% chance of crossing the line on its own. Unlike repeated looks, these segments are disjoint. No session sits in two of them. So the multiplication that was wrong above is roughly right here:
chance that at least one of twelve independent tests crosses = 1 - 0.95^12
= 1 - 0.5404
= 45.96%Simulating these exact twelve segments, at their real sizes, with no true effect anywhere, gives 45.6%. That is close to the multiplication, because disjoint segments really are nearly independent.
So on a test where the change did nothing at all, you would expect to find a winning segment in almost half of all such slices. Finding one is not surprising. Finding none would have been mildly surprising.
The standard correction is to divide the threshold by the number of comparisons:
0.05 / 12 = 0.0042The winning row is at 0.035, which is eight times larger. It does not survive.
And the winner's curse from earlier applies here in its purest form. The row was picked BECAUSE it was extreme. So its +71.4% is not an estimate of anything. It is the largest of twelve draws.
What a segment result IS good for
It is a hypothesis, and hypotheses are valuable. Ellinghay has just been handed a specific, testable idea it did not have on Monday, with a mechanism behind it: the fit panel helps shoppers who arrive cold from an advertisement. The only thing it may not do is spend money on it as though it were a finding.
The confirmation is a new test, declared in advance, on that segment alone, with the sample size worked out first. That is the same arithmetic as lesson 1. Run it on the segment's own base rate of 2.02% and its own traffic of 1,140 sessions an arm over four weeks. That is 285 an arm a week.
| Effect to be confirmed | Sessions an arm | Weeks at 285 an arm a week |
|---|---|---|
| The observed +71.4% | 2,081 | 7 |
| A more plausible +40% | 5,686 | 20 |
| A useful but modest +30% | 9,698 | 34 |
Seven weeks is affordable, and that is the test to run. If the effect really is anything like what the slice claimed, seven weeks will show it. If it comes back flat, Ellinghay has learned that the +71.4% was the largest of twelve draws. That is what the arithmetic already suggested, and it is worth seven weeks of one channel's traffic to establish before rebuilding the mobile product page around it.
Check yourselfA team runs a test for three weeks, checking the dashboard every morning, and ships on day 11 when the variant is up 22% at p = 0.04. What is wrong, and what would you say the change is actually worth?Show the answer
Three things, and they compound. Checking every morning for eleven days means eleven chances at the 5% threshold rather than one, so the false-positive rate for that decision is nowhere near 5%. The daily-peeking simulation in this lesson puts it above a quarter for a four-week test. Next, stopping on the first crossing means the estimate is conditional on being extreme, so +22% is biased upward whether or not there is a real effect underneath it. When the true effect is zero, daily peeking produces a median claimed lift of +33.3%. Third, the p of 0.04 was never a 5% bet in the first place, so it cannot be read at face value. What the change is worth is unknown. The honest next step is a fixed-sample test with the end date and the threshold declared before it starts. If the effect really is +22%, that test is short and cheap, and the arithmetic in lesson 1 will tell you exactly how short from your own base rate.
Prompt · Write the test plan before the split goes live
Before a test starts, when the temptation is to launch it and decide later.
Act as an experiment reviewer who has seen many tests read badly. I want to test [describe the change] on [describe the surface]. My figures: base conversion rate [rate]; sessions a day on this surface [number]; cost of the change [amount]; gross margin an order [amount]. Produce a one-page plan with exactly these headings, and refuse to leave any of them vague: - Primary metric, with its numerator and denominator stated. - Unit of analysis: session, visitor or order, and why. - Smallest effect worth acting on, derived from the cost and the margin rather than from what would be impressive. - Sample an arm and the end date that follow from it, with the arithmetic shown. - Looks: how many, when, and at what adjusted threshold, or none. - Guardrail metrics and what would make me stop for harm. - Segments I am allowed to report, named in advance and counted, with the threshold divided by that count. Then tell me plainly whether this test is worth running at all. If the traffic cannot detect the effect that matters, say so and suggest what to measure instead.
AI can make mistakes — check anything you act on.