Skip to lesson
Exit
Routing, failover and backpressure1 / 2

1 min lesson

Find the source of tail latency

Use percentiles and request dimensions to find the hop responsible for tail latency.

Step 1 of 2

The job description calls for a high-throughput, low-latency inference path. A product can have an acceptable average and still feel slow when its p99 is high.

Suppose p50 is 40 ms and p99 is 600 ms. Most requests are fast, but one in a hundred arrives much later. Inspect percentiles by product area, provider and request-path hop instead of relying on one average. Compare enough requests over the same time window and separate changes in traffic mix from changes in service time. Split the tail by model, request type, region and provider. Then inspect traces from the slow group to find the hop that consumed the extra time. Check sample counts and confidence intervals before treating a percentile change as a regression.

Learn more

Advanced table

Match each tail cause to a fix

Where the tail comes frommatch each cause to a fix

Source of tail
Cold connections
Why it spikes p99
The first request pays for a TLS and TCP handshake
Mitigation
Pool connections, use keep-alive and warm the pools
Source of tail
Runtime pauses
Why it spikes p99
A garbage collection pause delays one request
Mitigation
Reduce allocations on the path and tune the runtime
Source of tail
Head-of-line blocking
Why it spikes p99
A slow request stalls work behind it
Mitigation
Isolate requests and use bounded concurrency
Source of tail
Queueing delay
Why it spikes p99
The request waits before work starts
Mitigation
Shed early and size queues to the latency target
Source of tail
Slow provider
Why it spikes p99
One upstream raises end-to-end tail latency
Mitigation
Consider a hedge at a measured threshold when the gain justifies the cost

Measure the cause before choosing the fix.