1 min lesson
Trace each failure branch
Name the response and stopping condition for provider errors, overload, context limits and no viable route.
Step 1 of 2
The failure branches
Name the response for each failure and the point where the system stops retrying.
- For a temporary 5xx or timeout, spend from the retry budget and route to a healthy approved fallback.
- When admission is over budget, return an overload response with a backoff hint instead of queueing past the latency target.
- For context that exceeds a model limit, route to an approved model with enough context or reject before making the provider call.
- When no approved provider is viable, stop retrying and return a clear error instead of silently using an unsuitable model.
Admission checks the local request and token budgets. Routing chooses an approved healthy provider. The adapter serializes the shared request and applies the product timeout. Token deltas stream back while the trace records latency and cost. A temporary provider failure may spend from the retry budget and use an approved fallback. An open circuit skips that provider. A cancelled client request cancels the remaining provider work.
Move one request through admission, routing, the adapter, the provider and the stream. Then repeat the trace with a provider failure and overload. This makes dependencies, retry limits and exit conditions visible.