A/B Test Design & Statistical Rigor Questions
Designing and statistically defending a controlled online experiment: framing a testable hypothesis, defining control and treatment variants, choosing the randomization unit, setting the primary success metric, and computing sample size, power, and minimum detectable effect. Covers the statistical foundations that make a readout trustworthy, including hypothesis testing, p-values, confidence intervals, statistical vs practical significance, and Type I/II error. Emphasizes avoiding the common pitfalls that invalidate a test, such as peeking, multiple-comparison inflation, underpowered designs, and how test duration and stopping rules affect the validity of conclusions.
You run the same experiment across many countries, or across many device types and new-versus-returning users, and see a small but statistically significant uplift overall. Describe how you would assess whether the effect is genuinely heterogeneous across these segments: which interaction tests or models you would use, how you would power the per-segment analysis, and how you would correct for testing many segments at once so you don't just find noise. Compare full pooling, no pooling per segment, and partial pooling using a hierarchical model that borrows strength across segments, and recommend a rollout strategy given what you find.
Sample Answer
Direct answer
Assess heterogeneity with a formal treatment-by-segment interaction test rather than by eyeballing which country's point estimate looks different, and be honest that most individual segments are underpowered relative to the overall test. Because testing many segments at once inflates the odds that one looks significant by chance alone, correct for that multiplicity (a false-discovery-rate procedure for routine screening, Bonferroni's family-wise control when a single false positive would be costly) before trusting any one segment's result, and choose among full pooling (one overall effect), no pooling (an independent estimate per segment), and partial pooling (a hierarchical model that shrinks each segment's estimate toward the overall mean in proportion to how much data that segment actually has); partial pooling is the right production default for exactly this kind of many-segment, uneven-traffic setup, because it avoids both the false confidence of no pooling and the false uniformity of full pooling.
Structured elaboration
Testing for heterogeneity
Fit outcome ~ treatment + segment + treatment:segment, using both device type and region as covariates jointly rather than testing one dimension at a time, since a device-level pattern and a region-level pattern can be confounded with each other if only one is modeled. (In this regression-formula shorthand, ~ means "model the outcome using the terms on the right," and treatment:segment is the interaction term, the piece that tests whether the treatment effect itself varies by segment rather than just shifting the segment's baseline.) Test the joint significance of the interaction terms with a Wald or likelihood-ratio test, or a permutation test as a distribution-free alternative (any of the three answers the same question by a different route, so which one you pick matters less than actually pre-specifying and running one), to get one answer to "is there heterogeneity at all" before looking at any individual segment.
Why per-segment power is the first thing to check
A segment with a fraction of the overall traffic has a much larger minimum detectable effect (MDE) than the pooled test, using the same relationship between sample size and detectable effect as any power calculation:
MDE≈(z1−α/2+z1−β)n2p(1−p)
For an overall test with 300,000 users per arm and a 5% baseline conversion rate, at α=0.05 two-sided and 80% power:
MDEoverall≈(1.9600+0.8416)300,0002×0.05×0.95=2.8016×0.000563≈0.0016 (0.16pp,about 3.2% relative)
A single small geo with 3% of that traffic, 9,000 users per arm, has:
MDEsmall geo≈2.80169,0002×0.05×0.95≈0.0091 (0.91pp,about 18.2% relative)
That is roughly 5.8 times larger (matching 300,000/9,000), meaning that small geo's own data can only reliably detect an effect nearly six times bigger than what the overall test is powered for. A "no significant effect" reading for that geo on its own data is very often just this, not evidence the true effect is zero there.
Correcting for testing many segments at once
Underpowered segments are one problem; testing many of them at the same time creates a second, separate problem. Each segment's interaction test is its own hypothesis test, so scanning fifteen countries or five device types for "which one differs" is exactly the multiple-comparisons setting: even if the treatment effect is truly identical everywhere, running enough segment-level tests makes at least one spuriously significant result likely by chance alone. Two standard corrections apply once the segment set is pre-specified (decided before looking at the results, not assembled from whichever segments already look interesting):
- Bonferroni divides the significance threshold by the number of segments tested (roughly α/k for k segments). It controls the family-wise error rate (FWER), the probability of even a single false positive across the whole set, which makes it the conservative option: reach for it when acting on one wrongly-flagged segment would be expensive, for example a permanent per-country rollout split that is costly to unwind.
- Benjamini-Hochberg (BH) controls the false discovery rate (FDR) instead: the expected proportion of false positives among the segments you flag, not the chance of any false positive at all. It is less conservative than Bonferroni and is the standard choice for segment screening, because the usual goal when scanning many countries or device types is to shortlist which ones deserve a closer look, not to make one irreversible call per segment; tolerating a small, known rate of false leads among the flagged segments is the right trade for not burying the real ones.
Hierarchical shrinkage, covered in the pooling comparison below, answers the same over-testing question through a different mechanism rather than competing with these corrections. Bonferroni and BH both operate on the significance threshold a segment's p-value has to clear before you call it real; partial pooling never runs a per-segment significance decision in the first place; it pulls every segment's raw estimate toward the global mean in proportion to how little data that segment has, so a small segment cannot produce an extreme, attention-grabbing estimate purely from its own noise. A production pipeline commonly uses both together: an FDR-corrected interaction test to decide which segments are worth calling out as genuinely different, and hierarchical shrinkage on the point estimates actually shown to stakeholders, so the reported numbers are already regularized rather than raw.
Full pooling, no pooling, and partial pooling
| Approach | What it estimates | Bias | Variance | When it's the right call |
|---|---|---|---|---|
| Full pooling | One overall effect applied to every segment | Biased if heterogeneity is real, since it ignores it by construction | Lowest, because it uses all the data at once | The interaction test does not reject homogeneity, or segments are too small to estimate independently at all |
| No pooling | An independent effect estimate per segment | Unbiased in expectation, per segment | Highest, especially for small segments, as the MDE gap above shows | Only when each segment individually has enough traffic to be well-powered on its own |
| Partial pooling (hierarchical) | A segment-level effect modeled as drawn from a shared distribution across segments, shrinking each estimate toward the overall mean | Small, controlled bias traded for a large variance reduction on small segments | Between the two extremes, and adaptive: large, confident segments keep most of their own signal | The typical case: many segments of very uneven size, which is exactly a geography or device breakdown |
The shrinkage a hierarchical model applies has a standard closed form: a segment's raw estimate θ^i with sampling variance σi2 is pulled toward the grand mean θˉ by:
λi=τ2+σi2τ2,θ^ishrunk=λiθ^i+(1−λi)θˉ
where τ2 is the estimated between-segment variance.
To make this concrete, reuse the same two segments from the MDE comparison above: the overall test's 300,000-users-per-arm scale and the small geo's 9,000-users-per-arm scale, both at the 5% baseline conversion rate already used there. Their sampling variances follow directly from the same standard-error term used in the MDE formula, expressed here in percentage points:
σlarge=300,0002×0.05×0.95≈0.0563pp,σlarge2≈0.00317 pp2
σsmall=9,0002×0.05×0.95≈0.3249pp,σsmall2≈0.1056 pp2
Take an illustrative between-segment variance of τ2=0.09 pp2 (true segment effects varying by roughly ±0.3pp around the grand mean is a reasonable order of magnitude here: smaller than the small geo's own sampling noise, larger than the overall test's). The shrinkage weights are:
λlarge=0.09+0.003170.09≈0.966,λsmall=0.09+0.10560.09≈0.460
Suppose the pooled (grand-mean) estimate across all segments is θˉ=0.20pp, the large segment's own raw estimate happens to be θ^large=0.25pp, and the small geo's own raw estimate happens to be θ^small=0.85pp, more than four times the grand mean at face value. Shrinking each toward θˉ:
θ^largeshrunk=0.966×0.25+0.034×0.20≈0.248pp
θ^smallshrunk=0.460×0.85+0.540×0.20≈0.499pp
The large segment's raw estimate barely moves, 0.25pp to 0.248pp, because its own sampling variance is tiny next to τ2, keeping λlarge close to 1. The small geo's raw estimate moves a great deal, from 0.85pp down to about 0.50pp, more than half of its apparent excess over the grand mean pulled away, because its sampling variance is large next to τ2 and λsmall sits closer to 0. A striking small-segment number is exactly the case shrinkage is built to discount.
A weakly informative prior centered on the pooled estimate (for a Bayesian implementation) plays the same role as τ2 in a frequentist multilevel model: it is what lets a small geo borrow strength from the rest of the segments instead of standing entirely on its own thin data. A segment with a small σi2 (lots of its own data) keeps most of its own estimate; a segment with a large σi2 (little data) gets pulled hard toward the overall mean. This is the mechanism, not full pooling's blunt "ignore the segment" or no pooling's blunt "trust the segment's noisy number completely."
What to report to stakeholders
For each segment, report the shrunk point estimate, its credible or confidence interval, and something indicating how much shrinkage was applied (effective sample size, or the raw versus shrunk estimate side by side), alongside the single global interaction-test result and which multiplicity correction it used. Explain the shrinkage in plain language for a non-technical audience: "we pulled small countries' estimates toward the overall average because they don't have enough of their own data to stand alone yet," rather than presenting a raw per-country number that a stakeholder might otherwise take at face value.
Recommended rollout strategy
If the global interaction test does not reject homogeneity and the shrunk per-segment estimates are consistent in sign and magnitude with the pooled effect, roll out globally on the strength of the pooled result. If a segment's shrunk estimate remains meaningfully different in sign or magnitude even after shrinkage has pulled it toward the mean, and it survives the multiplicity correction above, that segment is a real candidate for differentiated treatment, but confirm it with a segment-targeted follow-up before committing production logic to a permanent split, the same discipline that applies to any post-hoc-adjacent segment finding.
Trade-offs & pitfalls
- Treating "not significant per segment" as "no effect in that segment." The MDE gap above is the mechanism; a small segment's null result is usually a power problem, not a finding.
- No pooling on small segments overstates confidence in noise. Reporting fifteen independent country-level point estimates without acknowledging their individual uncertainty invites a stakeholder to chase the noisiest ones.
- Full pooling by default hides real, actionable heterogeneity. If the interaction test does reject homogeneity, defaulting to one global number throws away a finding that could inform a real rollout decision.
- Explaining shrinkage badly. If a hierarchical model's output is presented as a black box, a segment owner who sees their raw number "corrected" downward without explanation will reasonably distrust the whole analysis.
- Skipping the multiplicity correction because shrinkage is already in use. Shrinkage regularizes the estimates; it does not by itself control how many segments get flagged as different across the whole set. Treating hierarchical modeling as a substitute for an FDR or family-wise correction, instead of a complementary tool, is how a many-segment scan quietly turns into a fishing expedition again.
A product team is designing an experiment that changes the homepage layout and needs to decide the unit of randomization: user id, session id, cookie, device, or household. For each candidate unit, describe the trade-offs (bias, cross-unit contamination, measurement noise) and explain how hash-based deterministic bucketing works in practice, including operational pitfalls such as changing hashing keys or salts mid-experiment. Recommend how you would detect and correct unit-mismatch problems after the experiment has run.
Sample Answer
Direct answer
The randomization unit should be the largest identity that is (a) stable over the experiment window and (b) matches the unit at which you will measure and report the outcome. For a homepage layout change with user-scoped conversion metrics, that is almost always user id when you have reliable logged-in identity; fall back to device id for logged-out mobile traffic, and treat cookie and session id as fallback-only units because they leak identity across the very boundary you are trying to hold fixed. The mechanism that turns "unit" into an actual bucket assignment is deterministic hash-based bucketing, and its main operational failure mode is touching the hash inputs (the salt or key) mid-experiment. Before any of that, though, you have to define who is even eligible to be in the experiment at all.
Structured elaboration
Defining the eligible population before choosing a unit
Unit choice is a second-order question; the first-order question is which units are even eligible to enter the experiment. For a mobile-only feature (say, a redesign shipped exclusively in the mobile app to a US audience), a desktop-only visitor cannot receive the treatment no matter which arm they land in, so randomizing across your full user base and then measuring outcomes at the account level silently dilutes the experiment: ineligible units get logged into both arms with a null "effect" (they cannot experience the change either way), which pulls the estimated treatment effect toward zero and inflates the sample size needed to detect a real one. The eligible population for a mobile-only US feature is the set of units that are (a) on the mobile platform that ships the feature, (b) in the targeted market (US), and (c) past whatever version or capability gate the feature requires; everyone outside that eligible population should be excluded from the experiment entirely, not folded into control by default. This is a distinct failure mode from picking the wrong unit: a design can choose a perfectly good unit (user id) and still be broken if a third of the "users" randomized into it were structurally incapable of ever seeing the treatment, whether the unit ultimately chosen within that eligible population is user, device, or session id.
Trade-offs by candidate unit
| Unit | Bias risk | Cross-unit contamination | Measurement noise | When it fits |
|---|---|---|---|---|
| User id | Low, if identity is stable and logged-in coverage is high | Low: one identity, one assignment across devices/sessions | Low: outcome aggregates cleanly to the assignment unit | User-scoped metrics (conversion per user, retention) with strong login coverage |
| Device id | Moderate: a shared household device mixes two people's behavior | Moderate: a device is stable, but a person moving across devices is not held fixed | Moderate | Logged-out or app-only surfaces where device is the closest stable identity |
| Cookie | Moderate to high: cleared on privacy sweeps, differs per browser | High: the same person can carry two cookies (two browsers) or none (private mode), landing in both arms or neither | High: undercounts multi-device, overcounts churny cookie population | Legacy web-only experiments with no login signal, used with caveats |
| Session id | High | High: the same user gets reassigned every new session, so the "treatment" a user experiences is not stable | High: session-level noise dominates any user-level signal | Only for genuinely session-scoped questions (e.g., a single-session UI micro-test) |
| Household | Low for spillover, but a distinct effective-sample-size cost | Low: contains treatment inside the family unit when family members influence each other's behavior | High variance per unit relative to user-level randomization, because you have fewer households than users | Shared-consumption products (streaming, shared carts) where one member's exposure changes another's behavior |
The two axes that matter are: does this unit stay attached to one treatment condition for the life of the experiment, and does it match the level at which you will later compute the metric. Session-level randomization on a homepage layout change fails both: a returning user can see version A on Monday and version B on Wednesday, so "the effect of the layout" is not well defined for that person, and if you then report a user-level conversion rate you are averaging over users who experienced a mix of both conditions.
Target-segment and control-group selection for a personalization test
Personalization experiments add a further wrinkle on top of eligibility and unit choice: because the treatment itself varies per person (each user's personalized experience differs from every other user's), you have to be explicit about two more things: which segment of the eligible population the test targets, and what the control group actually receives. A common setup: the target segment is the subset of eligible users with enough interaction history for the personalization model to act on (say, users with a minimum number of prior sessions); users below that threshold cannot be meaningfully personalized and should either be excluded from the test or routed to a defined fallback, rather than silently folded into a "control" group that has nothing to do with the personalization decision being tested. The control group, correspondingly, should receive a clearly defined non-personalized baseline (a fixed default ranking or layout), not "whatever the legacy system happened to show," so the measured effect is attributable to personalization itself rather than to incidental differences between the two code paths. Get target-segment or control-group definition wrong (an ill-specified segment boundary, or a control group that partially overlaps with treatment logic) and the measured lift reflects a spurious selection effect rather than the personalization algorithm's real value, no matter how correctly the underlying randomization unit and hash mechanism were implemented.
How hash-based deterministic bucketing works
In practice you do not store a per-user assignment row for every experiment. Instead you compute
bucket(u)=hash(u∥salt)modN
where u is the chosen unit id (user id, device id, etc.), the salt is a string unique to this experiment (often the experiment name or id), and N is the number of buckets (commonly 100 or 1000 for fine-grained traffic allocation). Buckets are then mapped to arms, e.g. buckets 0-49 to control and 50-99 to treatment for a 50/50 split. Because the hash is deterministic, the same unit id always lands in the same bucket for the same salt, which is what makes the assignment reproducible without a lookup table, and salting per-experiment is what makes assignment to experiment A independent of assignment to experiment B (so the same user can be validly in many concurrent, non-interacting experiments).
Operational pitfalls
- Changing the salt or hashing key mid-experiment. This is the single most common self-inflicted wound. It re-shuffles every unit into a new bucket, silently reassigning some fraction of users from control to treatment (or the reverse) partway through. The experiment now mixes users with a clean single-arm history and users who were exposed to both arms, which is exactly the session-level contamination problem from the table above, except it is invisible unless you log assignment history.
- Reusing a salt across experiments. If two unrelated experiments accidentally share a salt (or one is a substring of the identifier used in the other), their bucket assignments become correlated instead of independent, which breaks the assumption that concurrent experiments do not interfere with each other.
- Changing N or the bucket-to-arm mapping. Even without touching the salt, resizing the traffic split mid-flight (e.g., ramping from 5% to 50%) moves units across the arm boundary unless the mapping is designed to be monotonic (new traffic is added to existing arms rather than everyone being rehashed).
- Identity churn. A user id that gets merged, deleted, or re-issued (account merge, logout/login cycles that mint a new anonymous id) effectively becomes a new hash input mid-experiment, which has the same effect as a salt change for that user.
A finer-grained alternative: per-impression randomization
Every unit above is a person-shaped identity. Some teams instead randomize at the impression level, assigning a fresh coin flip to each page view or ranking request rather than to a person. This is occasionally used for high-frequency, low-persistence decisions (e.g., which of several ranking variants to serve on a given request) where you explicitly do not want a stable per-user experience. It is a different trade entirely from the table above: it eliminates any notion of "this user's assigned arm" (so it cannot answer a question about a durable, user-perceived change like a homepage layout), and it introduces strong intra-user correlation in the outcome data, since one person's many impressions are not independent draws, which inflates the effective variance if you naively treat impressions as independent observations in the analysis. Per-impression randomization is the right tool only when the thing being tested is meant to vary within a single user's experience; for a homepage layout, where the goal is to measure how a stable person-level experience changes behavior, it is the wrong granularity.
Detecting and correcting unit-mismatch after the fact
- Assignment-churn audit. From the exposure logs, compute the fraction of units that were logged under more than one arm during the experiment window. A near-zero rate is expected; anything material indicates contamination.
- Pre-period balance check. Compare the two arms on metrics measured before the experiment started (metrics that could not possibly be affected by treatment). An imbalance signals a broken randomization, not a broken hash necessarily, but it is the same diagnostic.
- Sample ratio mismatch check on the realized split, i.e., does the observed 50/50 (or intended ratio) actually hold at the analysis unit. A skew is a strong signal that the bucketing pipeline itself misbehaved.
- Timeline reconstruction. If churn is found, check the deployment log for the experiment: a salt, key, or bucket-count change on a specific date will produce a visible step change in the churn-rate-by-day series.
- Correction paths, in order of preference. Analyze by first-observed assignment only (treat each unit's initial exposure as its assignment, i.e., an intention-to-treat style rule, and accept the resulting dilution of the effect estimate); if the break has a clean date, restrict the analysis window to the stable period before or after it; if contamination is pervasive, drop the experiment's results for the affected window and rerun rather than trying to model around a broken assignment mechanism, since any post hoc adjustment for a data-dependent unit-mismatch is itself a source of bias.
Worked example
Suppose an app-only feature was randomized by session id and you are asked to sanity-check it before trusting the readout. You pull exposure logs and count, per user, the distinct arms they were logged under: 92,000 users saw only control, 91,500 saw only treatment, and 6,500 saw both. Churn rate is 6,500/(92,000+91,500+6,500)≈3.4%. That is a directly computed, reproducible number from the logs, not an assumption, and a value that high on a homepage-layout test (where the same person plausibly returns within the experiment window) is enough on its own to recommend re-running at user-id granularity rather than trying to salvage the session-level readout.
Trade-offs and pitfalls
- Choosing the "purest" unit (household) is not free: fewer independent units means higher variance per unit, so the same absolute effect needs more households than it would need users to reach the same precision. Unit choice is a bias-versus-noise trade, not a pure bias fix.
- A cookie- or device-based fallback is a compromise you should name explicitly to stakeholders, not a silent substitute for user id; report the estimated multi-device contamination rate alongside the headline result.
- An eligible population that is defined too loosely (e.g., randomizing all traffic instead of just the mobile-only, in-market segment) produces the same kind of diluted, biased-toward-zero readout as a bad unit choice, even when the unit itself is correct.
- Do not "fix" detected contamination by re-including the mixed-exposure users with a different weighting scheme chosen after seeing which way it moves the result; decide the exclusion or ITT rule before looking at the treatment effect.
What is an A/A test, and why would you run one before or alongside a real A/B test? Describe at least two valid use cases, such as validating the assignment and instrumentation pipeline or establishing a baseline-variance estimate, and two limitations or common misinterpretations of A/A testing. If an A/A test shows a statistically significant difference between the two identical groups, what steps would you take to root-cause it?
Sample Answer
Direct answer
An A/A test randomly splits traffic into two groups that receive the identical experience and compares their metrics as if they were a real experiment. You run one to validate the assignment and measurement pipeline before trusting a real A/B result on the same platform: since both groups get the same product, any statistically significant difference between them signals a problem in the pipeline (randomization, instrumentation, or analysis) rather than a real effect, because by construction there is no effect to detect.
Structured elaboration
Two valid use cases
- Validating the assignment and instrumentation pipeline. Confirms that the bucketing hash actually produces the intended split ratio, that each unit sees a stable, single experience, and that event logging correctly attributes actions to the assigned arm end to end (client instrumentation through to the analysis table).
- Establishing a baseline-variance estimate. Because there is no true effect, the spread of the A/A metric difference across many runs (or across a well-chosen resampling of the same data) tells you what "just noise" looks like for this metric on this population, which is useful input for planning: it is a sanity check on your variance assumptions, not a substitute for a proper power calculation.
What to check while it runs
- The realized split ratio against the intended one (a sample-ratio check): meaningfully off the intended ratio (say, 50/50 skewing to 49/51 in a way that recurs, not a single noisy day) points at a bucketing bug before you even look at outcome metrics.
- Core funnel and event counts by arm (sessions, page views, primary conversion event) to confirm the two arms are tracked with equal fidelity, not just equal traffic.
- Whether the metric of interest for the upcoming real experiment behaves as expected in the A/A read, since that is the metric whose baseline variance you actually need.
- Run it for at least one full natural cycle of the traffic (typically a full week, to span weekday/weekend mix) rather than a single day, since a one-day A/A window can look clean by luck or flagged by a day-specific anomaly that has nothing to do with the platform.
An A/A test is, at its core, a targeted way to surface three distinct failure classes: instrumentation errors (events not logged or misattributed), non-random assignment (the bucketing hash is not producing a genuinely random, independent split), and sampling biases (the two arms end up systematically different in composition despite a technically-random split, e.g., a bot-filtering rule that behaves differently by arm). Each class points at a different fix, which is why segmenting the flagged difference (below) matters more than the raw significance flag itself.
Two limitations or common misinterpretations
- A clean A/A result is not proof the pipeline is bug-free. With enough traffic, small true differences in a specific test can still slip through if the bug is intermittent (e.g., only affects a rare browser) or if the metric checked in the A/A test is not the one that will matter in the real experiment. Absence of a flagged difference is reassurance, not a guarantee.
- A single significant A/A result does not, by itself, mean the pipeline is broken. At a conventional significance threshold, some fraction of A/A tests will show a "significant" difference purely by chance even with a perfectly correct pipeline; treat one flagged metric as a prompt to investigate, not as an automatic verdict, especially if you are checking many metrics at once and did not correct for that.
Root-causing a significant A/A result
- Recheck the sample ratio first. A skewed split is the fastest, most common finding and points straight at a bucketing bug rather than a downstream measurement issue.
- Segment the difference. Break the flagged metric down by platform, geography, and new-vs-returning user; a difference concentrated in one segment (e.g., one app version) points at an instrumentation bug specific to that segment rather than a global randomization failure.
- Check for a known confound in how the two arms are served, such as one arm being disproportionately served through a code path with different latency or caching behavior, which is functionally a version-of-treatment bug even though no real treatment was intended.
- Re-run before escalating, if the first read used a short window: a single noisy day is a weaker signal than a difference that persists across multiple independent A/A windows.
- If it persists and is not explained by a segment or a known bug, treat the underlying real-experiment platform as unvalidated until the discrepancy is resolved; shipping A/B decisions on top of an unexplained A/A anomaly defeats the purpose of running the check at all.
Worked example
A team runs an A/A test ahead of a planned homepage experiment and flags a significant difference in click-through rate. Step 1, sample ratio: 50.1% vs 49.9%, within normal noise, so not a bucketing problem. Step 2, segmentation: the CTR gap is near zero on Android and web but noticeably present on iOS. Step 3: engineering finds one arm's iOS client is on an older app version with a slightly different default tab order, an artifact of how the A/A test's client-side flag was staged rather than anything about the experiment platform itself. The root cause is a version-of-treatment bug traced to a real, checkable fact (the iOS staging config), not a p-value alone; the fix is correcting the staged rollout, not adjusting the metric.
Trade-offs and pitfalls
- Running A/A tests constantly, on every metric, invites exactly the false-alarm problem described above; use them at meaningful checkpoints (new platform, new metric pipeline, post-incident) rather than as a standing tax on every experiment.
- Do not use a single A/A run's variance estimate as your only power-planning input if you have a more direct historical baseline available; treat it as a cross-check.
- A quiet A/A test on a low-traffic metric provides much weaker reassurance than the same result on a high-traffic metric, because a real problem of a given size is harder to detect with less data; do not treat "clean" as equally strong evidence across metrics of very different volume.
Users increasingly interact with a product across multiple devices and login states, which creates duplicate identities: for example, web experiment assignment is cookie-based while the mobile app uses a device id, and after backend identity merging many users turn out to have been placed into both variants. Explain how cross-device identity resolution and deduplication affect experiment assignment and analysis, and propose practical strategies to minimize the bias from duplicate counting and cross-variant contamination.
Sample Answer
Direct answer
When assignment happens per-device (a cookie on web, a device id on the app) but the real unit of interest is the person, users who touch the product on multiple devices get assigned independently on each device, so some of them land in both control and treatment at once. That breaks the assumption that each experimental unit receives exactly one arm: it dilutes the measured treatment effect, since a "contaminated" user's behavior is influenced by both arms, and it can double count outcomes if the same person's actions are logged and analyzed once per device-identity rather than once per person. The fix is to randomize and log at the most persistent identity you actually have, then dedupe and correct for the identities you had to resolve after the fact rather than at assignment time.
Structured elaboration
Where the bias enters
- Assignment stage: a user with a web cookie and a mobile device id gets two independent coin flips. If both land the same way, no harm; if they split, that person is exposed to treatment and control simultaneously, a violation of SUTVA (the assumption that one unit's outcome does not depend on another instance of its own assignment).
- Analysis stage: if the analysis unit is "device" rather than "resolved person," a contaminated person's activity appears once in each arm's numerator, and the two rows are not independent observations even though the analysis code treats them as such, which understates the true variance.
- Coverage bias: identity resolution itself makes mistakes (false merges linking two different people, false splits treating one person as two); if those errors are not random with respect to treatment, they introduce their own bias on top of the contamination.
Practical strategies
- Randomize at the most stable identity you have. Prefer a logged-in account id over a device id or cookie whenever the user is authenticated; fall back to a deterministic device hash only for logged-out traffic, and treat that population as a separate, lower-confidence stratum in reporting.
- Log everything needed to resolve identity after the fact. Persist device id, cookie id, and account id (hashed, respecting privacy) on every assignment and every outcome event, even when the assignment itself was made at device level, so contamination can be measured and corrected during analysis rather than discovered too late.
- Define the primary analysis on the resolved (canonical) identity, not the raw assignment record: after backend identity merge, collapse a contaminated user into a single row and apply an explicit, pre-registered rule for what arm they count as, for example "any-device-treatment counts as treated," reported alongside a stricter "single-device-only" rule as a sensitivity check.
- Quantify and bound the bias rather than ignore it. Report the primary result plus at least two sensitivity analyses: one restricted to users seen on exactly one device (removes contamination but shrinks the sample), and one using the any-device-treatment rule (keeps the full sample but is a diluted estimate of the true per-exposure effect).
- Use cluster-robust standard errors at the resolved-identity level so that outcomes from the same person are not treated as independent observations even after collapsing to one row per person, since a person can still contribute multiple sessions or events.
Worked example
Assume 20% of users are active on exactly two devices (web and app) and 80% are single-device (an illustrative, stated split). Assignment happens independently per device with probability 0.5 to treatment. For a two-device user, the four equally likely device-pair outcomes are (control, control), (control, treatment), (treatment, control), (treatment, treatment), each with probability 0.25:
P(both control)=P(both treatment)=0.25,P(split, i.e. contaminated)=0.5
So among the 20% of users who are two-device, half get split across arms, which is 10% of the total user base. If a contaminated user's outcome is counted in both the treatment and control totals rather than resolved to one arm, then 10% of the treatment-arm numerator and 10% of the control-arm numerator are contributed by the exact same set of people, which both understates the between-arm difference and violates the independence assumption behind the standard error calculation. Restricting the primary analysis to the 90% of users who are single-device or resolve cleanly to one arm removes the contamination at the cost of 10% of the sample, which should be reflected directly in the power calculation for the corrected analysis.
Trade-offs and pitfalls
- The "any-device-treatment" rule is conservative and interpretable but structurally dilutes the estimated effect toward zero for contaminated users, since they experienced a mix of both arms; don't present it as an unbiased estimate of the pure per-exposure effect.
- Restricting to single-device users is cleaner statistically but changes who the estimate applies to: if multi-device users differ systematically (often more engaged, higher-value), the single-device estimate may not generalize to the full user base.
- Identity resolution is itself a model with false-merge and false-split error rates; a resolution pipeline retrained or changed mid-experiment can shift the contamination rate over time and should be monitored, not assumed constant.
- Logging every identifier needed for later resolution has real privacy and storage cost; agree on hashing and retention policy with the privacy function before instrumenting, not after a contamination investigation is already underway.
You are evaluating a price increase (for example, raising a marketplace take rate or introducing a new fee) in a two-sided marketplace with network effects between buyers and sellers. Design an experiment that accounts for spillovers between the two sides: specify the randomization scheme, including whether to randomize by buyer, seller, or a shared cluster, how you would detect and quantify cross-side externalities, and what analysis approach you would use to estimate the long-run revenue impact under these network effects.
Sample Answer
Direct answer
The key design choice is randomizing at a unit large enough to contain the cross-side spillover a price or fee change creates: typically a market cluster (a geography, city, or self-contained supply region) rather than individual buyers or individual sellers, because a marketplace's two sides interact through the same local supply-demand pool. Within that, you deliberately vary treatment intensity across clusters (a partial-saturation or dose-response design) so you can separate the direct pricing effect from the cross-side externality it triggers, and you plan the revenue read as a staged measurement: an early operational window for the direct effect, followed by a longer holdout-based window because churn and re-equilibration on a two-sided market take longer to surface than a single-user metric would.
Structured elaboration
Why buyer-only or seller-only randomization fails here
If you randomize individual sellers into a higher take rate, a treated seller may raise prices, list less, or churn; buyers who would have transacted with that seller instead transact with a control-arm seller in the same market. The buyer side is now indirectly treated regardless of which arm assigned it, which is the same SUTVA-style interference problem as a social feed, except the shared medium is the local marketplace instead of a social graph. Randomizing individual buyers has the mirror-image problem on the seller side. Either choice contaminates the arm you intended to leave clean.
Randomization scheme
- Unit: market cluster (geography, city, or another boundary where most matching happens locally, e.g., delivery radius). This keeps most buyer-seller matching internal to a single treatment condition.
- Design shape: partial saturation. Instead of a flat 50/50 split, assign clusters to a small number of treatment intensities (for example, no increase, a modest increase, a larger increase) rather than a single on/off arm. This lets you trace how the cross-side response scales with the size of the change, which a single treatment level cannot distinguish from a fixed step change.
- Randomize at the cluster level, stratified on baseline liquidity (existing buyer-to-seller ratio, transaction volume) so clusters that already look structurally different are balanced across intensities before you start, reducing the chance that a treatment-intensity effect is confounded with a pre-existing market difference.
Detecting and quantifying the cross-side externality
- Because clusters are internally exposed to one intensity, you can compare a directly-treated side's metric (e.g., seller take-rate exposure) against the other side's metric within the same cluster (buyer conversion, buyer price sensitivity) to see if the fee change on sellers moved buyer-side behavior, and by how much, as intensity increases.
- The partial-saturation design turns this into a dose-response check: plot the buyer-side metric against assigned intensity across clusters. A flat line across intensities is evidence of a contained direct effect; a sloped line is direct, in-cluster measurement of the spillover, not an assumption about its existence.
- Compare a cluster's realized outcome to adjacent, untreated clusters it plausibly shares supply with (e.g., neighboring cities where sellers can relist), which is a direct check for leakage across the cluster boundary itself, not just across sides within a cluster.
Estimating the long-run revenue impact
- Short-run direct revenue (immediate take-rate math: transactions times the new fee) is mechanical and available immediately, but it is not the number that matters, because it ignores behavioral response.
- The number that matters is the net of three components measured over a longer window: the direct fee revenue, minus revenue lost to seller churn or delisting, minus revenue lost to buyer-side friction from any resulting price or availability change. Each of these three needs its own time horizon: fee revenue is immediate, seller churn plays out over weeks as sellers decide whether to stay, and buyer-side effects play out over the buyer's own return cadence.
- Because of that lag structure, hold a subset of clusters as a long-run holdout past the point where you make the initial ship decision. This is what lets you catch delayed seller attrition or buyer defection that would not have shown up in an early readout, and it is a standard practice for any monetization change with a plausible slow-churn tail, not something specific to marketplaces.
Worked example
A delivery marketplace tests a take-rate increase across 40 city clusters, split into two treatment intensities plus control (roughly 13-14 clusters each, stratified on baseline order volume so the three groups start with comparable liquidity). At the direct level, transaction-weighted take-rate revenue rises with intensity, as expected mechanically. The diagnostic step is checking buyer-side order volume within the same clusters: if buyer order volume also declines with intensity (a negative slope across the three intensity levels, measured, not assumed), that decline is the in-cluster, dose-response evidence of the cross-side externality: sellers responded to the higher take rate by raising prices or delisting, and buyers responded to that. The ship decision then nets the two effects (higher unit take rate, lower volume) into an actual revenue trajectory rather than trusting the mechanical fee-revenue number alone.
Trade-offs and pitfalls
- Cluster randomization costs statistical power relative to individual-level randomization, because the effective sample size is the number of clusters, which is typically far smaller than the number of users; this needs a longer test or fewer, larger clusters, and it is a real cost you should state up front rather than discover after the fact.
- A short observation window will understate the true cost of the change, because seller churn and buyer defection both lag the price change; shipping on the early direct-revenue number alone is the single most common mistake in this design.
- Neighboring-cluster leakage (a seller in a treated city relisting in an adjacent control city) is a real risk for marketplaces with mobile supply; check it explicitly rather than assuming cluster boundaries are airtight.
- Resist reaching for a full structural or instrumental-variable model as the default; those are appropriate when randomization is genuinely unavailable, but the partial-saturation cluster design above gives a directly measured effect and should be preferred whenever you can actually randomize.
Unlock Full Question Bank
Get access to all 22 A/B Test Design & Statistical Rigor interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.