Skip to lesson
Exit
Routing, failover and backpressure1 / 2

1 min lesson

Measure the inference path

Use inference-specific metrics and useful dimensions to distinguish routing, provider and capacity faults.

Step 1 of 2

Add inference-specific metrics

Latency, traffic, errors and saturation show general service health. Add TTFT, inter-token latency, failover, shedding, token use and provider outcomes to explain what happened on the inference path. Split each metric by product area, model, provider, region and outcome so one failing route does not disappear inside a global average. Alert on sustained movement away from a measured baseline, not any nonzero failover or shed event. Keep enough labels to locate a fault, but control label cardinality so the metrics system remains usable during an incident. Pair an alert with the view needed to decide whether to route away, reduce admission or roll back. Page only when there is a defined response.

Learn more

Advanced table

Read the inference metrics

Metric
TTFT
What it tells you
Time to first token, the latency users feel
Alert when
p99 breaches the product area's budget
Metric
Inter-token latency, ITL
What it tells you
Stream smoothness during generation
Alert when
ITL spikes. The provider is degrading mid-stream
Metric
Cache hit rate
What it tells you
Prefix and response cache effectiveness
Alert when
Drops sharply. Cost and latency regress
Metric
Failover rate
What it tells you
How often requests leave the preferred route
Alert when
Rises outside a planned rollout. Check health, policy and quality
Metric
Shed rate
What it tells you
How much load you reject
Alert when
Rises above the expected baseline or stays elevated
Metric
Per-provider error rate
What it tells you
Which upstream returns failures
Alert when
One provider moves above its error baseline

Failover rate and shed rate can reveal routing or capacity pressure before overall availability moves.

Interview practice

I would set separate availability, p99 latency and cost targets for each product area. I would manage them with an error budget and pause risky routing changes when it is exhausted. Failover rate shows traffic leaving the primary provider. Shed rate shows active rejection. Both can move before a coarse availability number changes.