1 min lesson
Protect waiting users during overload
Prioritize waiting users during contention while preserving bounded capacity for background work.
Step 1 of 2
Priority and QoS classes
Interactive completions have a tighter latency target than background indexing or batch evaluations. Under contention, protect work with a person waiting and defer background jobs first.
- Class
- Interactive / latency-critical
- Example surface
- Tab, ⌘K
- Under contention
- Reserve capacity, shed last and never queue past budget
- Class
- Interactive / tolerant
- Example surface
- Agent, chat
- Under contention
- Served, but may degrade to a fallback model
- Class
- Bulk / background
- Example surface
- Re-index, batch eval
- Under contention
- Defer first or apply a low rate limit
| Class | Example surface | Under contention |
|---|---|---|
| Interactive / latency-critical | Tab, ⌘K | Reserve capacity, shed last and never queue past budget |
| Interactive / tolerant | Agent, chat | Served, but may degrade to a fallback model |
| Bulk / background | Re-index, batch eval | Defer first or apply a low rate limit |
Reserve capacity for each class, use fair queuing within a class and give waiting users priority during contention. Tab does not wait behind a batch job, and background work retains a bounded share.
A request that cannot finish inside its latency target should not wait in a longer queue. Rejecting it returns capacity and lets the client back off. Letting it time out wastes work and may add a retry to the original load.
Return a distinct overload response with a backoff hint. If a client retries every generic 500 immediately, shedding can amplify the spike instead of reducing it.