1 min lesson
Choose fallbacks by product need
Choose product-specific fallbacks and state when the gateway must stop retrying.
Step 1 of 2
Quality-aware fallback
A fallback model may differ in speed, cost, context length or quality. Choose acceptable fallbacks for each product area and record when the system uses them. Do not assume the same tradeoff works for a short completion and a long agent task.
I would define fallbacks for each product area. A short completion may tolerate a smaller model if quality measurements support it. A long agent task should prefer a model in the same quality tier or return a clear retryable error. I would track fallback rate and quality so reduced service is visible.
If every approved fallback is unavailable, stop retrying. Shed work that no longer fits its latency budget and return a clear retryable error. State this limit instead of claiming failover can absorb every outage.
Learn more
Optional practice
Set a retry budget
QWhy is a global retry budget safer than a fixed per-request retry count when a provider has a broad outage?