A/B Test Significance Calculator
A green 'winner' badge in your testing tool is the start of the question, not the answer. Enter sessions and orders for each version to see the p-value, the plausible range of the lift, and whether the test had enough traffic to detect a lift that matters.
95% interval for the difference: +0.05 to +0.67 pts. This test could reliably detect lifts of about 17.9% or more. Checked it early?
Talk to us about your measurementControl converts 480 of 20,000 sessions (2.4%); the variant converts 552 of 20,000 (2.76%). That is a 15% relative lift with p ≈ 0.023, significant at 95%, with a 95% interval of about +0.05 to +0.67 percentage points.
Read the interval, not just the p-value
The calculator runs a two-sided two-proportion z-test, the standard test behind most A/B tools. The p-value tells you how surprising the gap would be if the variant did nothing. The confidence interval tells you what you actually need for a business case: the range of lifts consistent with the data.
Plan on the low end of that interval. Winners picked because they crossed a significance line overstate their effect on average, so the lift you see after rollout is usually smaller than the lift in the test.
A non-significant test is not proof of no effect
The calculator also shows the smallest relative lift your sample could reliably detect. If that number is 25% and you were hoping for 8%, the test never had a chance. You learned that the change is not enormous, nothing more. Size the next test before launch with the sample size calculator.
Frequently asked questions
What confidence level should I use?
95% is the standard. Use 99% when a false winner would trigger a large, hard-to-reverse investment such as a full site redesign or a pricing change.
Can I test revenue per session with this?
No. This calculator handles conversion rates. Revenue per session is skewed by a small number of large orders and needs a t-test or bootstrap on per-session revenue.
Why did my winner not hold up after launch?
Usually some mix of peeking, testing many variants or metrics and reporting the best, and the natural overstatement of effects selected for significance. Fix the sample in advance and read the result once.
More free tools
- Break-Even ROASThe ROAS and CPA your paid media must clear before a single order makes money.
- Sample SizeHow many sessions and how many weeks your test needs before you launch it.
- Stop or ContinueA straight answer on whether your test can end, and what checking it daily has cost you.
- SRM CheckerCatch a broken traffic split before a corrupted test drives a real decision.
Ready to prove your marketing ROI?
Book a free 30-minute consultation. Bring your ad account numbers; leave knowing which channels earn their budget.
Free. No commitment.