1 min lesson
Route around a provider outage
Filter eligible routes, choose a healthy provider and contain a failure without removing unrelated routes.
Step 1 of 2
The current job description names a clear goal: no single provider outage causes user-visible degradation. A good design routes around a failing provider before most requests reach it.
Use several controls because they address different failures. Health routing avoids a provider that is already unhealthy. A circuit breaker stops new calls after repeated failures. A bounded retry can recover a brief error. A fallback list chooses another provider when the first one is unavailable. Health is only one routing condition. First filter routes by model capability, data policy, region, context limit and available capacity. Then choose among the eligible healthy routes. Track health by provider, model and region so a failure on one route does not remove every route from that provider.
Learn more
Full explanation
Compare provider failure controls
Interactive diagram. Tab through its regions; each focused region shows its detail in the panel below.
The order shows how broad an outage each layer can absorb. Earlier layers prevent failures that later ones can only recover from.