1 min lesson
Limit gateway overhead
Measure and reduce gateway overhead without shifting latency into contention or buffering.
Step 1 of 2
Limit gateway overhead
Measure how much latency the gateway adds before the provider starts generating. Keep connection setup, local queueing and synchronous work out of the request path where possible. Measure gateway processing separately from provider time to first token. Give local work its own overhead target and test it under representative concurrency. A change that lowers average CPU but adds lock contention can still make p99 worse, so compare the full latency distribution before and after the change.
- Pool connections and keep them warm so no hot request pays for a handshake.
- Use asynchronous, non-blocking I/O so one slow upstream does not hold a worker needed by other requests.
- Keep required parsing and hashing small. Defer nonessential accounting and avoid synchronous work that blocks other requests.
- Forward the stream as it arrives. Buffering the full response delays the first token until generation finishes.
Learn more
Full explanation
Set a measured hedge threshold
Choose a hedge threshold from measurements
A hedge can reduce a slow provider tail, but each hedge starts more work. Set the trigger from the live latency distribution and the product area's cost and latency targets.
- Start with a measured percentile
- The current p95 can be a starting point, but test whether it meets the product area's latency target.
- Recompute the threshold
- Provider health and traffic change the distribution, so review the threshold continuously.
- Measure duplicate cost
- Track extra request starts and tokens. A useful latency gain still needs an acceptable cost.
- Cancel the other attempt
- Cancel the slower request as soon as the winning response is safe to serve.