# Experiment Plan — Reason-Matched Cancellation

**Status:** Pre-registration draft  
**Decision owner:** Product  
**Analysis owner:** Data Science  
**Data:** Synthetic sizing inputs for portfolio demonstration

## Decision to make

Should we replace the current cancellation path with a reason-aware, trust-first flow for direct web subscribers?

## Hypothesis

For eligible direct web subscribers who start cancellation, showing clear consequences and at most one reason-matched alternative will increase **voluntary 30-day retained value** by at least 3.0 percentage points without reducing cancellation completion, increasing support contact, or lowering trust.

## Variants

**Control:** Existing self-serve cancellation flow.

**Treatment:**

1. optional reason selection;
2. consequence summary;
3. at most one eligible reason-matched alternative;
4. equally prominent cancellation action; and
5. durable confirmation.

The treatment is evaluated as a bundle for the first test. A follow-up test may isolate consequence clarity from the alternative.

## Population and assignment

- **Unit of randomization:** billing account, not session or user.
- **Eligibility:** active direct web subscriber; cancellation initiated from account settings.
- **Exclusions:** app-store billing, enterprise/multi-seat, collections, refund dispute, prior experiment exposure, internal/test accounts.
- **Persistence:** assignment survives repeat starts for the full experiment.
- **Planned exposure:** 50/50 after staged safety ramp.

## Metrics

### Primary

**Voluntary 30-day retained value rate**

Eligible cancellation starts where the account accepts an alternative, remains paid or intentionally paused 30 days later, and completes at least one meaningful product action within the observation window ÷ eligible cancellation starts.

### Guardrails

| Metric | Rule |
|---|---|
| Cancellation completion | Non-inferior within −1.5 pp and absolute ≥ 90% |
| Cancellation-related support contacts / 1,000 starts | Must not increase by > 5% |
| Post-flow trust score | Non-inferior within −0.1 on five-point scale |
| Billing/cancellation incident rate | No statistically or operationally material increase |
| Page performance | p75 additional load time ≤ 300 ms |

### Diagnostic

- reason capture and skip rate;
- alternative view and acceptance;
- time to complete;
- backtracks;
- meaningful action by reason and alternative;
- re-cancellation at day 7, 30, and 60;
- contribution margin after discount/pause cost.

## Decision rules

| Outcome | Decision |
|---|---|
| Primary ≥ +3.0 pp and every guardrail passes | Ramp to 25%; inspect segment consistency |
| Primary positive but < +3.0 pp; guardrails pass | Iterate eligibility or response; do not broad-ramp |
| Primary passes; any trust guardrail fails | Stop and treat as loss |
| Primary neutral; contacts materially decline | Consider shipping clarity-only version |
| Effect concentrated in one reason/plan | Limit rollout to supported segment; run confirmatory test |

No decision will be made on the first seven days because retained value requires a 30-day window.

## Sizing inputs

The minimum detectable effect, baseline, traffic, and power calculation are [REDACTED — commercial analytics]. Before launch, Data Science will attach the executable sizing notebook and record:

- baseline and lookback window;
- expected eligible starts per week;
- MDE and rationale;
- alpha and power;
- expected observation + maturation period; and
- adjustment for repeated looks, if any.

## Quality checks before reading lift

1. Verify experiment assignment before UI render.
2. Check sample-ratio mismatch overall and by billing channel.
3. Confirm event order: start → response → decision → completion.
4. Compare missingness across variants.
5. Validate duplicate account suppression.
6. Confirm no concurrent pricing or billing experiment contaminates exposure.

## Stopping rules

Stop exposure immediately for:

- confirmed duplicate charge or cancellation state mismatch;
- cancellation completion below 85% over any 24-hour window with adequate volume;
- sensitive reason or free text appearing in experiment payloads;
- support contacts > 20% above control for two consecutive daily reads; or
- material accessibility failure in the treatment path.

Do not stop early because the primary metric crosses significance once.

## Heterogeneity plan

Pre-specified cuts:

- reason;
- monthly vs annual plan;
- tenure band;
- prior meaningful use; and
- market/locale.

Segment reads are directional unless independently powered. They may narrow a rollout; they will not justify a broad claim.

## Risks to interpretation

- A 30-day pause may shift churn beyond the observation window.
- Users who would have stayed may cannibalize into a cheaper option.
- Trust survey response may be non-random.
- Reason selection may reflect offer-seeking rather than true motivation.
- Novelty can temporarily increase engagement.

The final readout must include day-60 re-cancellation and contribution margin before declaring durable value.

