Skip to lesson
Exit
Routing, failover and backpressure1 / 2

1 min lesson

Trace one request through the design

Trace one request from admission to streaming, then repeat it through overload and provider failure.

Step 1 of 2

Use one request to explain how the proposed design works. Trace it from admission to the first streamed token, then repeat the trace with a provider failure and an overloaded gateway.

A request trace shows how the components interact and where each control applies. It also reveals what happens when no provider or queue can accept more work. Start with the request's product area, logical id, required capabilities, token estimate and deadline. At each stage, name the input, the decision and the output passed forward. Reuse those same inputs when tracing a 5xx, overload or cancellation so the changed branch and stopping condition are clear. End with the metrics that confirm the branch behaved as intended and stayed inside its retry and cost budgets.

Learn more

Full explanation

Follow the request path

One request from admission to stream

Interactive diagram. Step through it with the Next and Previous controls below, or Tab to a region to read its detail.

diagram: flow

Each stage owns one class of failure. That keeps responsibilities from overlapping.

Learn more

Advanced table

Compare short and long request settings

Where each technique plugs in

Stage
Admission
Technique
Rate limit and load shed
Failure it prevents
A spike reaching workers and providers unchecked
Stage
Routing
Technique
Health, policy and optional cache affinity
Failure it prevents
A request reaching an unhealthy or disallowed provider
Stage
Gateway
Technique
Product timeout and measured hedge threshold
Failure it prevents
One slow provider consuming the latency budget
Stage
Provider call
Technique
Circuit breaker
Failure it prevents
Repeated calls to a failing upstream
Stage
Retry path
Technique
Budgeted retry and approved fallback
Failure it prevents
A single provider failure ending every request

Each stage owns a specific decision and failure response.

Use different settings for short and long requests

In this proposed design, short interactive completions and long agent tasks share components but use different timeouts, hedge policies, fallbacks and cancellation rules.

Tab request

Shorter prompt with a tight interactive latency target.

Use a short timeout and consider hedging only if measurements justify it.

Approve a smaller fallback only when quality checks support the tradeoff.

Agent request

Longer, more expensive and more likely to be cancelled during several steps.

Avoid duplicate execution unless the latency gain justifies the cost.

Prefer an equivalent-quality fallback and propagate cancellation to the provider.