1 min lesson
Run a controlled online experiment
Set the outcome, guardrails, audience, duration and stop conditions before an online test begins.
Step 1 of 2
Confirm with a controlled online experiment
Online experiments show whether a change helps developers in real sessions. Limit the initial audience according to risk, keep treatment assignment stable and define stop conditions before traffic moves.
- Primary outcome
- Choose the user or code outcome before launch and state how it will be calculated.
- Guardrails
- Track latency, token cost, tool errors, cancellations and other harms the change could introduce.
- Ramp
- Start with an audience small enough to limit harm, then expand only while outcome and guardrail checks pass.
- Duration and uncertainty
- Run for a preplanned window that covers relevant usage cycles and report confidence intervals or another suitable uncertainty measure.
Do not choose the winning metric after reading the result or stop the test as soon as a favorable difference appears. Log exclusions, sample-ratio checks, guardrail breaches and rollout changes so the decision can be reproduced.
Describe the full decision path. State the hypothesis, offline task set, primary online outcome, guardrails, initial audience, stop conditions and readout. Explain which result would block a rollout even if another metric improved.