A/B Test Design & Statistical Rigor Questions

Designing and statistically defending a controlled online experiment: framing a testable hypothesis, defining control and treatment variants, choosing the randomization unit, setting the primary success metric, and computing sample size, power, and minimum detectable effect. Covers the statistical foundations that make a readout trustworthy, including hypothesis testing, p-values, confidence intervals, statistical vs practical significance, and Type I/II error. Emphasizes avoiding the common pitfalls that invalidate a test, such as peeking, multiple-comparison inflation, underpowered designs, and how test duration and stopping rules affect the validity of conclusions.

HardTechnical
41 practiced

You run the same experiment across many countries, or across many device types and new-versus-returning users, and see a small but statistically significant uplift overall. Describe how you would assess whether the effect is genuinely heterogeneous across these segments: which interaction tests or models you would use, how you would power the per-segment analysis, and how you would correct for testing many segments at once so you don't just find noise. Compare full pooling, no pooling per segment, and partial pooling using a hierarchical model that borrows strength across segments, and recommend a rollout strategy given what you find.

MediumTechnical
78 practiced

A product team is designing an experiment that changes the homepage layout and needs to decide the unit of randomization: user id, session id, cookie, device, or household. For each candidate unit, describe the trade-offs (bias, cross-unit contamination, measurement noise) and explain how hash-based deterministic bucketing works in practice, including operational pitfalls such as changing hashing keys or salts mid-experiment. Recommend how you would detect and correct unit-mismatch problems after the experiment has run.

EasyTechnical
38 practiced

What is an A/A test, and why would you run one before or alongside a real A/B test? Describe at least two valid use cases, such as validating the assignment and instrumentation pipeline or establishing a baseline-variance estimate, and two limitations or common misinterpretations of A/A testing. If an A/A test shows a statistically significant difference between the two identical groups, what steps would you take to root-cause it?

HardTechnical
43 practiced

Users increasingly interact with a product across multiple devices and login states, which creates duplicate identities: for example, web experiment assignment is cookie-based while the mobile app uses a device id, and after backend identity merging many users turn out to have been placed into both variants. Explain how cross-device identity resolution and deduplication affect experiment assignment and analysis, and propose practical strategies to minimize the bias from duplicate counting and cross-variant contamination.

HardTechnical
45 practiced

You are evaluating a price increase (for example, raising a marketplace take rate or introducing a new fee) in a two-sided marketplace with network effects between buyers and sellers. Design an experiment that accounts for spillovers between the two sides: specify the randomization scheme, including whether to randomize by buyer, seller, or a shared cluster, how you would detect and quantify cross-side externalities, and what analysis approach you would use to estimate the long-run revenue impact under these network effects.

Unlock Full Question Bank

Get access to all 22 A/B Test Design & Statistical Rigor interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.