Skip to lesson
Exit
Routing, failover and backpressure1 / 2

1 min lesson

Trace every provider attempt

Keep every retry, hedge and fallback in one logical trace while measuring each attempt separately.

Step 1 of 2

Tracing across providers

A trace should show whether delay came from admission, routing, the adapter, the provider or the return stream. Keep retries, hedges and fallbacks under one logical trace so duplicate attempts remain connected. Give every upstream attempt its own child span and attempt id while retaining one logical request id. This separates provider time from retry delay without counting one user action as several requests. Sample ordinary successes if volume requires it, but retain errors, fallbacks and slow-tail traces for diagnosis.

  • Propagate a trace and idempotency id from admission through every retry, hedge and fallback so duplicated work is still one logical trace.
  • Give the routing decision, adapter serialization, provider TTFT and stream duration separate spans so the trace identifies the slow hop.
  • Tag spans with the provider, model, product area and outcome. Record served, failed over or shed so you can split tail latency by any field.
Learn more

Full explanation

Roll out a model with a tested exit

Safe rollout of a new model or provider

A new model can change quality, latency and cost. Limit the first live audience and make removal from routing a configuration change with a measured propagation time.

  1. 1Mirror approved traffic without serving the new model's output. Compare latency, cost and quality while still accounting for privacy and provider spend.
  2. 2Route a small live share and hold it while reviewing product-area SLOs, quality measurements and failover rate.
  3. 3Use a configuration control to remove the model from routing without a new code deployment. Test how long that change takes to reach every region.
  4. 4Increase traffic only while the error budget and quality checks remain healthy. Stop the rollout when either one regresses.
Limit the first audience

Scope a canary to one region, product area and provider when policy allows it. A routing control can then remove the change before a regression reaches more traffic. Keep data residency and other placement rules ahead of latency or cost optimization.

Learn more

Optional practice

Find a spreading failure

QWhich two inference-specific metrics can show that a provider or capacity problem is spreading before overall availability falls?