# Experiment Readout — Confidence Gate v1

**Decision:** Ramp to 50% for billing/policy and 25% for account access; hold technical support at 10%  
**Status:** Decision complete  
**Experiment window:** 21 days + 7-day quality review  
**Owners:** Product, Support Operations, Applied ML, Data Science  
**Data classification:** Synthetic portfolio simulation

> Integrity note: This readout uses synthetic aggregate data. It is not a claim of production impact. Counts, rates, confidence intervals, and p-values are internally consistent and included to demonstrate decision quality.

## 1. Executive readout

Confidence Gate reduced critical policy errors from **2.31% to 1.37%**:

- absolute effect: **−0.94 percentage points**;
- relative effect: **−40.7%**;
- 95% confidence interval: **−1.47 to −0.42 pp**;
- two-sided `p = 0.00045`.

All pre-registered operational guardrails passed. Median handle time increased 4.6%, below the +8% limit. Agent throughput declined 3.3%, below the −5% limit. Escalation rose 0.5 pp, below the +1.0 pp limit. CSAT was directionally positive.

The aggregate result is strong enough to ship, but not strong enough to ignore heterogeneity. Billing/policy drove most of the benefit; technical support was directionally positive but inconclusive.

### Decision

1. Ramp billing/policy queues from 25% to 50%.
2. Ramp account-access queues from 10% to 25% with weekly error audit.
3. Hold technical support at 10% while improving technical-document coverage.
4. Keep auto-send out of scope.
5. Do not claim a general 41% safety improvement across all support work.

## 2. Why we ran this test

The existing copilot generated fluent replies but did not distinguish:

- a supported answer from a plausible one;
- a low-risk instruction from a financial or privacy promise; or
- model uncertainty from missing or conflicting evidence.

Agents could open the knowledge base, inspect the account, and verify a draft, but that work happened in separate tools. Review behavior was inconsistent, especially under queue pressure.

Confidence Gate placed three signals inside the draft workflow:

1. claim-level evidence coverage;
2. risk tier based on claim type; and
3. confidence calibrated on held-out support cases.

The product required verification or escalation when the policy engine detected a high-risk claim, missing support, conflicting sources, or confidence below the risk-tier threshold.

## 3. Hypothesis and decision criteria

### Hypothesis

For eligible AI-assisted support conversations, adding a targeted evidence gate will reduce critical policy errors by at least 30% relative without increasing median handle time by more than 8%, escalation by more than 1.0 pp, or reducing agent throughput by more than 5%.

### Pre-registered criteria

| Metric | Ship criterion | Observed | Result |
|---|---:|---:|---|
| Critical policy error rate | ≥ 30% relative reduction | 40.7% reduction | Pass |
| Median handle time | ≤ +8% | +4.6% | Pass |
| Escalation rate | ≤ +1.0 pp | +0.5 pp | Pass |
| Agent throughput | ≥ −5% | −3.3% | Pass |
| CSAT | No material decline | +0.03 / 5 | Pass |

The primary metric required both statistical evidence and an operationally meaningful effect. Guardrails were decision gates, not secondary decoration.

## 4. Design

| Element | Choice | Rationale |
|---|---|---|
| Unit of randomization | Agent-week | Prevents agents switching review behavior between variants in one shift |
| Population | AI-assisted, English-language, eligible queues | Aligns exposure with supported evidence system |
| Control | Draft + existing knowledge links | Current agent workflow |
| Treatment | Draft + claim evidence + risk/confidence gate + escalation action | Tests the complete decision support loop |
| Duration | 21 days | Covers weekday/weekend mix and repeat agent exposure |
| Quality maturation | 7 days | Allows audit sample completion and late corrections |
| Primary analysis | Intent to treat | Preserves randomization |

### Exclusions

- conversations not eligible for AI drafting;
- unsupported languages;
- active incident queues;
- legal hold, threat, or abuse workflows already requiring specialist handling;
- bot-only conversations; and
- internal/test accounts.

## 5. Validity checks

### Sample ratio

Expected assignment was 50/50. Observed:

| Variant | Conversations | Share |
|---|---:|---:|
| Control | 5,012 | 49.93% |
| Confidence Gate | 5,027 | 50.07% |

No sample-ratio mismatch was detected. Assignment coverage and missing-event rates were comparable across variants.

### Pre-exposure balance

The variants were directionally balanced on queue mix, agent tenure, language eligibility, prior weekly throughput, and historical quality score. No standardized difference exceeded the pre-set review threshold.

### Instrumentation

- 99.4% of treatment drafts carried a gate decision event.
- 98.9% of completed conversations had a final send or escalation outcome.
- Audit assignment occurred after experiment assignment and before the auditor saw the agent action.
- Auditor disagreement was resolved blind to variant.

## 6. Results

### Primary metric — critical policy error

A critical error is a sent response containing an unsupported or incorrect financial, privacy, security, eligibility, or account-access claim that could cause customer harm or require formal correction.

| Variant | Critical errors | Eligible conversations | Rate |
|---|---:|---:|---:|
| Control | 116 | 5,012 | 2.31% |
| Confidence Gate | 69 | 5,027 | 1.37% |

Treatment minus control: **−0.94 pp**; relative change: **−40.7%**; 95% CI: **−1.47 to −0.42 pp**; two-sided `p = 0.00045`.

This clears the 30% relative-reduction threshold.

### Guardrails

| Metric | Control | Gate | Difference | Limit | Result |
|---|---:|---:|---:|---:|---|
| Median handle time | 8.7 min | 9.1 min | +4.6% | ≤ +8% | Pass |
| Escalation rate | 8.4% | 8.9% | +0.5 pp | ≤ +1.0 pp | Pass |
| Agent throughput | 42.1 / day | 40.7 / day | −3.3% | ≥ −5% | Pass |
| CSAT | 4.42 / 5 | 4.45 / 5 | +0.03 | No decline | Pass |

Median handle time hides a longer tail. p90 handle time increased 7.2%. This remains inside the operational threshold but will be monitored during ramp.

### Diagnostic behavior

| Behavior | Control | Gate | Read |
|---|---:|---:|---|
| Agent edited draft before send | 31.2% | 24.9% | Evidence may improve first-draft usefulness |
| Knowledge source opened | 18.6% | 43.7% | Verification behavior increased |
| Unsupported high-risk claim sent | 1.41% | 0.62% | Main mechanism appears active |
| Gate overridden | — | 6.8% | Requires continued audit |

These are mechanisms, not success metrics. A lower edit rate is useful only when audited quality improves.

## 7. Segment read

| Segment | Control | Gate | Absolute | Relative | Decision |
|---|---:|---:|---:|---:|---|
| Billing / policy | 3.71% (67/1,804) | 1.82% (33/1,810) | −1.89 pp | −50.9% | Ramp to 50% |
| Technical | 1.39% (25/1,800) | 1.10% (20/1,810) | −0.28 pp | −20.5% | Hold at 10% |
| Account access | 1.70% (24/1,408) | 1.14% (16/1,407) | −0.57 pp | −33.3% | Ramp to 25% |

The experiment was powered for the aggregate, not every segment. Segment estimates are used to bound rollout, not to make independent causal claims.

### Interpretation

Billing/policy combines high baseline risk with strong structured evidence. The gate can identify promises and retrieve a narrow policy set.

Technical support has lower baseline critical-error incidence and broader evidence. The same gate adds workflow cost while catching fewer severe errors. The next product problem is evidence coverage, not a lower confidence threshold.

## 8. What surprised us

**Confidence was not the dominant trigger.** Most prevented errors involved evidence conflict or a high-risk claim with no supporting source—even when model confidence was above 85%.

That changes the roadmap. Better calibration matters, but the near-term moat is claim-level provenance plus risk-aware policy, not a universal confidence score.

## 9. Decision and rollout

### Ship

- claim-level evidence for all eligible queues;
- hard verification gate for billing/policy high-risk claims;
- escalation path for conflicting sources;
- audit event stream and override reasons.

### Hold

- mandatory gate across every technical answer;
- any auto-send behavior;
- agent performance scoring based on overrides;
- customer-visible confidence labels.

### Ramp plan

| Queue | Current | Next | Review gate |
|---|---:|---:|---|
| Billing / policy | 25% | 50% | 7 days; error and p90 handle time |
| Account access | 10% | 25% | 14 days; quality audit ≥ 300 cases |
| Technical | 10% | 10% | Improve evidence recall before new test |

Full operational detail: [ROLLOUT_PLAN.md](ROLLOUT_PLAN.md).

## 10. Limitations

1. The test covers English-language, AI-eligible conversations only.
2. Auditor labels encode policy interpretation and are not perfectly objective.
3. Randomization by agent-week reduces contamination but can leave cluster correlation; final inference should use cluster-robust errors.
4. The 21-day window does not measure long-term agent deskilling or reliance.
5. Support queue mix may shift during incidents or policy changes.
6. CSAT response is selective and underrepresents silent harm.
7. The experiment does not establish readiness for autonomous sending.

## 11. What would reverse this decision

Pause or roll back if any of the following occur during ramp:

- critical error rate returns above 2.0% in a ramped queue;
- p90 handle time increases more than 10% for three consecutive days;
- override-related errors exceed 0.25% of gated conversations;
- evidence freshness incidents produce an incorrect policy answer; or
- agent-reported trust in the tool falls by 0.3 or more on the five-point pulse.

## 12. Next experiments

1. **Technical evidence retrieval:** improve source recall before retesting the gate.
2. **Gate copy:** compare reason-specific instructions with generic “verify” language.
3. **Override capture:** test structured reason collection without punitive signaling.
4. **Evidence freshness:** test visible source date and owner on policy-sensitive claims.
5. **Deskilling audit:** compare unaided quality after repeated treatment exposure.

## 13. Bottom line

The experiment supports a bounded launch, not a victory lap. Confidence Gate materially reduces critical errors where risk and evidence are structured. It does not yet justify universal friction or autonomy.

