1 min lesson
Choose the lever that matches the delay
Use traces and request outcomes to choose the smallest change that addresses the measured delay.
Step 1 of 2
Choose a lever that matches the delay
Use traces and request outcomes to find where time and money are spent. Each lever below solves a different cause and should be kept only when an experiment improves the target metric without breaking a guardrail. Map each slow span to the smallest change that can remove it. If client work dominates, tune debounce and cancellation. If inference dominates, compare eligible model routes with the same requests. If equivalent requests repeat, test a cache key that covers code state, settings and policy. If network time varies by region, compare regional routing with capacity and data placement costs. Measure a baseline first, change one lever at a time and compare usefulness, tail latency and cost. Remove a lever when its operational burden exceeds the measured gain.
Learn more
Advanced table
Compare latency levers
- Lever
- Debounce and cancel
- What it buys
- Avoids work for input that changed before a result could be used
- Cost / catch
- A long delay feels unresponsive, while a short delay may still send wasteful requests
- Lever
- Smaller eligible model
- What it buys
- Reduces inference time and cost for requests it can handle
- Cost / catch
- Quality may fall, so route only where measured outcomes remain acceptable
- Lever
- Speculative decoding
- What it buys
- Can reduce the time before useful output is available
- Cost / catch
- Incorrect speculation spends compute and may add system complexity
- Lever
- Caching
- What it buys
- Reuses results for an equivalent context and request
- Cost / catch
- The key must include every input that can change correctness or policy
- Lever
- Prefetch during a pause
- What it buys
- Moves some work before the next explicit request
- Cost / catch
- Unused results still cost money and may become stale
- Lever
- Regional routing
- What it buys
- Can reduce network time for distant users
- Cost / catch
- Adds deployment, capacity and data-placement constraints
| Lever | What it buys | Cost / catch |
|---|---|---|
| Debounce and cancel | Avoids work for input that changed before a result could be used | A long delay feels unresponsive, while a short delay may still send wasteful requests |
| Smaller eligible model | Reduces inference time and cost for requests it can handle | Quality may fall, so route only where measured outcomes remain acceptable |
| Speculative decoding | Can reduce the time before useful output is available | Incorrect speculation spends compute and may add system complexity |
| Caching | Reuses results for an equivalent context and request | The key must include every input that can change correctness or policy |
| Prefetch during a pause | Moves some work before the next explicit request | Unused results still cost money and may become stale |
| Regional routing | Can reduce network time for distant users | Adds deployment, capacity and data-placement constraints |