A/B Test Analysis + the Peeking Problem

Full statistical treatment of a mobile game A/B test, including a simulation of how much "checking daily and stopping early" inflates false positives.

p = 0.074
retention_1 (not significant)
p = 0.0016
retention_7 (significant)
21.6%
false-positive rate from daily peeking
5.0%
restored after correction

The peeking problem, visualized

A hypothetical: 8,000 simulated trials under the null hypothesis (no real effect), tested repeatedly over 14 simulated days. Naive peeking, stopping as soon as p<0.05, roughly quadruples the real false-positive rate above the nominal 5%. A simulation-calibrated significance boundary brings it back down.

Chart showing false-positive rate climbing to 21.6% under naive daily peeking, vs. staying near 5% with a corrected significance boundary

The question

Cookie Cats moved its first progression gate from level 30 to level 40. Does that hurt player retention?

retention_1: gate_30 44.8% vs. gate_40 44.2%, diff 0.59pp, 95% CI [-0.06pp, 1.24pp], not significant
retention_7: gate_30 19.0% vs. gate_40 18.2%, diff 0.82pp, 95% CI [0.31pp, 1.33pp], significant, gate_40 is worse
Guardrail (game rounds played, outlier excluded): Mann-Whitney p = 0.051, borderline
Recommendation: don't ship the gate_40 change. No next-day benefit, and a real hit to 7-day retention.

Why this exists

Full write-up, including the power calculation, guardrail methodology, and what the peeking correction is actually doing, is in the repo README and the executable notebook.

View the repo on GitHub →