1 min lesson
Shape traffic before provider limits
Enforce provider request and token limits across gateway instances before upstream rejection.
Step 1 of 2
Shape traffic before the provider rejects it
Use each provider's known or measured request and token limits to shape traffic at the gateway. Waiting for upstream 429 responses can trigger retries after the system is already over capacity. Several gateway instances may admit traffic at once. Do not let each instance assume it owns the provider's full limit. Partition the budget between instances or reconcile a shared quota often enough to prevent over-admission. Keep a reserve for failover so a provider outage does not immediately exhaust the surviving route.
- 1Represent each provider's request and token limits with a local token bucket.
- 2Admit work against the bucket. Shed it or queue it briefly when it would exceed the available budget.
- 3Reserve capacity for bursts and failover traffic instead of running at the provider's stated maximum.
- 4If 429 responses rise despite the model, lower the local admission rate and review the limit data.