AI product · Trust · Support operations

Portfolio simulation: Synthetic conversations, policies, experiment inputs, and results.

An AI copilot that knows when not to send.

Confidence Gate inserts evidence and targeted human judgment between a fluent draft and a risky customer promise.

01 · Product experience

Put proof next to the promise.

Choose a case. The gate changes its intervention based on claim risk, evidence coverage, and model confidence.

Conversation

Annual plan refund

High risk
MK

I renewed yesterday but meant to cancel. Can you confirm that the full annual charge will be refunded today?

Account
Annual · renewed 18h ago
Region
British Columbia
Prior refund
None

AI draft

Response review

71% confidence
I’m sorry about the surprise renewal. I can confirm your annual charge will be fully refunded today, and access will end immediately.
1

“fully refunded today”No supporting policy found for this account state.

2

“access will end immediately”Conflicts with annual-plan access policy.

Gate decision

Verification required

A high-risk financial promise is unsupported. Verify policy or escalate before sending.

!

02 · Friction map

The dangerous gap is between “looks right” and “is supported.”

Open the complete map ↗
01

Triage

Understand intent

Context is split across ticket, account, and policy.
02

Generate

Draft arrives fluent

Polish raises perceived correctness.
03

Verify

Find the proof

Switching tools makes checking expensive.
04

Decide

Edit or escalate

Ownership is unclear in ambiguous cases.
05

Learn

Close the loop

Corrections rarely improve the system.

03 · Experiment readout

Safety improved. The rollout still needs a boundary.

Synthetic 21-day A/B test · 10,039 eligible conversations · randomized by agent-week.

Decision

Ramp selected queues to 50%

Ship for billing/policy and account-access work. Hold technical support for more evidence. Keep auto-send out of scope.

SHIP, BOUNDED

Primary metric

Critical policy error rate

Control
2.31%
Gate
1.37%
−0.94 pp

95% CI −1.47 to −0.42 pp · p < 0.001

Control n=5,012 · Gate n=5,027

All conversations and results are synthetic. Statistical values are calculated from the displayed counts.