1 min lesson
Budget latency end to end
Split the product latency target across each request-path hop and fix the measured bottleneck.
Step 1 of 2
Budget the milliseconds end to end
Budget time across the client-to-gateway hop, admission, routing, the provider call, generation and the return stream. Measure each hop so you know where a change can affect the product area's target.
When asked to make an interactive completion faster, begin with p95 and p99 by hop. Use warm pooled connections for handshake delay, bounded queues for admission delay and a measured hedge threshold for a slow provider tail. Choose the fix that matches the data.
Hedging too early starts duplicate work for ordinary requests and can raise latency by adding load. Limit it to the measured slow tail and track the extra request and token cost.
Learn more
Optional practice