Skip to lesson
Exit
Routing, failover and backpressure1 / 2

1 min lesson

Hedge only measured slow requests

Choose when to hedge from measured latency and cost, then account for every attempt.

Step 1 of 2

Use hedging for selected slow requests

A hedge starts a second request after a delay, uses the first acceptable response and cancels the other request. It can reduce tail latency, but it also starts duplicate provider work. Hedge only work that is safe to duplicate. Confirm that a second attempt has capacity and that cancellation reaches the provider. A cancelled attempt may still consume billed tokens, so compare the saved tail latency with the cost of every attempt. Disable the policy when extra load erases the latency gain.

When to hedge
Choose by product area
Consider hedging short, latency-sensitive requests. Avoid it for long requests unless the measured latency gain justifies the extra work.
Keep one logical result
Track both attempts under one request id, forward only the winner and cancel the other attempt promptly.
Learn more

Full explanation

Keep duplicate attempts idempotent

Keep retries and hedges idempotent

A retry or hedge can send the same logical request upstream twice. The gateway must return one result and account for the attempts without applying an effect twice.

  • Assign an idempotency key at admission before a retry or hedge can create another attempt.
  • Forward one winning response and drop later duplicates.
  • Make completions side-effect-free at the gateway layer so a duplicate is wasteful but never incorrect.
  • Record provider cost for every attempt, while applying the winning result and any user-facing usage only once.