Same data, three statistical views
Testvoy does not reduce every experiment to one shiny score. The same experiment data is read with Bayesian decision probability, Frequentist validation and SRM traffic quality checks. This helps reduce early-decision risk, allocation issues and explanation gaps across teams.
Bayesian: probability to beat control and expected loss
Frequentist: p-value and confidence interval
SRM: whether traffic allocation deviates from expectation
Revenue: lift can be read with revenue impact too
When should I decide?
A high lift on a small sample can look exciting while still being risky. Ending a test without reading minimum sample, traffic split, SRM, bot filtering and segment behavior together can produce the wrong winner.
- Is sample size sufficient?
- Is there an SRM warning?
- Any bot or dropped event signal?
- Does the CI support the business decision?
- Do revenue and segments point the same way?
Example report-reading flow
Example: Variant B looks strong in total conversion. Before deciding, read funnel steps, segment breakdown, revenue impact, bot-filtering signals and SRM warnings. This order reduces the risk of deciding early on one metric.
- 01
Read the Best variant card
- 02
Check Bayesian/Frequentist signals
- 03
Look for SRM warning
- 04
Inspect funnel drop-off steps
- 05
Compare device, country and UTM segments
- 06
Check whether revenue or RPV supports the same direction
- 07
Create a shareable report
Variant B shows +8.4% lift on signup_completed.
No SRM warning.
Mobile segment is stronger; desktop is neutral.
Checkout drop-off did not change.
Decision: move Variant B to 25% rollout.Practical scenario
A growth team changes CTA copy on the pricing page, connects signup_started and signup_completed goals, then reads Google Ads and returning visitor segments separately. If the result is promising, they share a report link with the client or leadership.
Common mistakes
Most mistakes come from wrong project keys, overly broad selectors, missing goals, staging/prod mixups, early decisions on small samples or skipping mobile QA.
Naming an experiment 'test'
Launching without a goal
Deciding on total CVR only
Sending screenshots instead of shareable reports