1 min lesson
Find the source of tail latency
Use percentiles and request dimensions to find the hop responsible for tail latency.
Step 1 of 2
The job description calls for a high-throughput, low-latency inference path. A product can have an acceptable average and still feel slow when its p99 is high.
Suppose p50 is 40 ms and p99 is 600 ms. Most requests are fast, but one in a hundred arrives much later. Inspect percentiles by product area, provider and request-path hop instead of relying on one average. Compare enough requests over the same time window and separate changes in traffic mix from changes in service time. Split the tail by model, request type, region and provider. Then inspect traces from the slow group to find the hop that consumed the extra time. Check sample counts and confidence intervals before treating a percentile change as a regression.
Learn more
Advanced table
Match each tail cause to a fix
Where the tail comes frommatch each cause to a fix
- Source of tail
- Cold connections
- Why it spikes p99
- The first request pays for a TLS and TCP handshake
- Mitigation
- Pool connections, use keep-alive and warm the pools
- Source of tail
- Runtime pauses
- Why it spikes p99
- A garbage collection pause delays one request
- Mitigation
- Reduce allocations on the path and tune the runtime
- Source of tail
- Head-of-line blocking
- Why it spikes p99
- A slow request stalls work behind it
- Mitigation
- Isolate requests and use bounded concurrency
- Source of tail
- Queueing delay
- Why it spikes p99
- The request waits before work starts
- Mitigation
- Shed early and size queues to the latency target
- Source of tail
- Slow provider
- Why it spikes p99
- One upstream raises end-to-end tail latency
- Mitigation
- Consider a hedge at a measured threshold when the gain justifies the cost
| Source of tail | Why it spikes p99 | Mitigation |
|---|---|---|
| Cold connections | The first request pays for a TLS and TCP handshake | Pool connections, use keep-alive and warm the pools |
| Runtime pauses | A garbage collection pause delays one request | Reduce allocations on the path and tune the runtime |
| Head-of-line blocking | A slow request stalls work behind it | Isolate requests and use bounded concurrency |
| Queueing delay | The request waits before work starts | Shed early and size queues to the latency target |
| Slow provider | One upstream raises end-to-end tail latency | Consider a hedge at a measured threshold when the gain justifies the cost |
Measure the cause before choosing the fix.