Free tool · Experimentation

A/B Test Significance Calculator

A green 'winner' badge in your testing tool is the start of the question, not the answer. Enter sessions and orders for each version to see the p-value, the plausible range of the lift, and whether the test had enough traffic to detect a lift that matters.

Result
Significant
p-value
0.0232
Relative lift
+15.00%
Control CVR
2.4%
Variant CVR
2.76%

95% interval for the difference: +0.05 to +0.67 pts. This test could reliably detect lifts of about 17.9% or more. Checked it early?

Talk to us about your measurement
Worked example

Control converts 480 of 20,000 sessions (2.4%); the variant converts 552 of 20,000 (2.76%). That is a 15% relative lift with p ≈ 0.023, significant at 95%, with a 95% interval of about +0.05 to +0.67 percentage points.

Read the interval, not just the p-value

The calculator runs a two-sided two-proportion z-test, the standard test behind most A/B tools. The p-value tells you how surprising the gap would be if the variant did nothing. The confidence interval tells you what you actually need for a business case: the range of lifts consistent with the data.

Plan on the low end of that interval. Winners picked because they crossed a significance line overstate their effect on average, so the lift you see after rollout is usually smaller than the lift in the test.

A non-significant test is not proof of no effect

The calculator also shows the smallest relative lift your sample could reliably detect. If that number is 25% and you were hoping for 8%, the test never had a chance. You learned that the change is not enormous, nothing more. Size the next test before launch with the sample size calculator.

Frequently asked questions

What confidence level should I use?

95% is the standard. Use 99% when a false winner would trigger a large, hard-to-reverse investment such as a full site redesign or a pricing change.

Can I test revenue per session with this?

No. This calculator handles conversion rates. Revenue per session is skewed by a small number of large orders and needs a t-test or bootstrap on per-session revenue.

Why did my winner not hold up after launch?

Usually some mix of peeking, testing many variants or metrics and reporting the best, and the natural overstatement of effects selected for significance. Fix the sample in advance and read the result once.

More free tools

Currently accepting new clients

Ready to prove your marketing ROI?

Book a free 30-minute consultation. Bring your ad account numbers; leave knowing which channels earn their budget.

Review your current attribution and measurement setup
Identify the highest-priority evidence gaps
Leave knowing which channels earn their budget

Free. No commitment.

No obligation. We reply within one business day.