1 min lesson
Trace one request through the design
Trace one request from admission to streaming, then repeat it through overload and provider failure.
Step 1 of 2
Use one request to explain how the proposed design works. Trace it from admission to the first streamed token, then repeat the trace with a provider failure and an overloaded gateway.
A request trace shows how the components interact and where each control applies. It also reveals what happens when no provider or queue can accept more work. Start with the request's product area, logical id, required capabilities, token estimate and deadline. At each stage, name the input, the decision and the output passed forward. Reuse those same inputs when tracing a 5xx, overload or cancellation so the changed branch and stopping condition are clear. End with the metrics that confirm the branch behaved as intended and stayed inside its retry and cost budgets.
Learn more
Full explanation
Follow the request path
Interactive diagram. Step through it with the Next and Previous controls below, or Tab to a region to read its detail.
Each stage owns one class of failure. That keeps responsibilities from overlapping.
Learn more
Advanced table
Compare short and long request settings
Where each technique plugs in
- Stage
- Admission
- Technique
- Rate limit and load shed
- Failure it prevents
- A spike reaching workers and providers unchecked
- Stage
- Routing
- Technique
- Health, policy and optional cache affinity
- Failure it prevents
- A request reaching an unhealthy or disallowed provider
- Stage
- Gateway
- Technique
- Product timeout and measured hedge threshold
- Failure it prevents
- One slow provider consuming the latency budget
- Stage
- Provider call
- Technique
- Circuit breaker
- Failure it prevents
- Repeated calls to a failing upstream
- Stage
- Retry path
- Technique
- Budgeted retry and approved fallback
- Failure it prevents
- A single provider failure ending every request
| Stage | Technique | Failure it prevents |
|---|---|---|
| Admission | Rate limit and load shed | A spike reaching workers and providers unchecked |
| Routing | Health, policy and optional cache affinity | A request reaching an unhealthy or disallowed provider |
| Gateway | Product timeout and measured hedge threshold | One slow provider consuming the latency budget |
| Provider call | Circuit breaker | Repeated calls to a failing upstream |
| Retry path | Budgeted retry and approved fallback | A single provider failure ending every request |
Each stage owns a specific decision and failure response.
Use different settings for short and long requests
In this proposed design, short interactive completions and long agent tasks share components but use different timeouts, hedge policies, fallbacks and cancellation rules.
Shorter prompt with a tight interactive latency target.
Use a short timeout and consider hedging only if measurements justify it.
Approve a smaller fallback only when quality checks support the tradeoff.
Longer, more expensive and more likely to be cancelled during several steps.
Avoid duplicate execution unless the latency gain justifies the cost.
Prefer an equivalent-quality fallback and propagate cancellation to the provider.