Skip to lesson
Exit
Capstone: Full Mock Loop & Self-Exam1 / 2

1 min lesson

Scenario: a provider's p99 spikes 3x at peak

Start at the first move in "Scenario: a provider's p99 spikes 3x at peak" and carry it through to the proof.

Step 1 of 2

Scenario: a provider's p99 spikes 3x at peakReason it out loud

  1. 1Measure before acting. Split TTFT from ITL. Rising TTFT points at queueing or prefill saturation upstream; rising ITL points at decode contention or the provider shrinking your batch share.
  2. 2Localize. Is it one provider or all of them? One region? Correlate with your QPS - did the spike come from us or them?
  3. 3Mitigate. Shift health-based routing weight off the slow provider, enable hedging for the latency-critical surfaces with the cost cap and let admission control shed low-value Tab to protect Agent.
  4. 4Verify. Watch p99 per surface recover and watch the hedge cost and failover counters so the cure isn't worse than the disease.
Say it like this

"First I'd separate TTFT from ITL, because they point at different causes. If TTFT is the spike, I'd suspect queueing and shift routing weight off that provider, then hedge the Tab traffic with a one-hedge cap. I'd confirm with per-surface p99 and keep an eye on hedge cost so I'm not paying double to fix a 5-minute blip."