1 min lesson
Measure the inference path
Use inference-specific metrics and useful dimensions to distinguish routing, provider and capacity faults.
Step 1 of 2
Add inference-specific metrics
Latency, traffic, errors and saturation show general service health. Add TTFT, inter-token latency, failover, shedding, token use and provider outcomes to explain what happened on the inference path. Split each metric by product area, model, provider, region and outcome so one failing route does not disappear inside a global average. Alert on sustained movement away from a measured baseline, not any nonzero failover or shed event. Keep enough labels to locate a fault, but control label cardinality so the metrics system remains usable during an incident. Pair an alert with the view needed to decide whether to route away, reduce admission or roll back. Page only when there is a defined response.
Learn more
Advanced table
Read the inference metrics
- Metric
- TTFT
- What it tells you
- Time to first token, the latency users feel
- Alert when
- p99 breaches the product area's budget
- Metric
- Inter-token latency, ITL
- What it tells you
- Stream smoothness during generation
- Alert when
- ITL spikes. The provider is degrading mid-stream
- Metric
- Cache hit rate
- What it tells you
- Prefix and response cache effectiveness
- Alert when
- Drops sharply. Cost and latency regress
- Metric
- Failover rate
- What it tells you
- How often requests leave the preferred route
- Alert when
- Rises outside a planned rollout. Check health, policy and quality
- Metric
- Shed rate
- What it tells you
- How much load you reject
- Alert when
- Rises above the expected baseline or stays elevated
- Metric
- Per-provider error rate
- What it tells you
- Which upstream returns failures
- Alert when
- One provider moves above its error baseline
| Metric | What it tells you | Alert when |
|---|---|---|
| TTFT | Time to first token, the latency users feel | p99 breaches the product area's budget |
| Inter-token latency, ITL | Stream smoothness during generation | ITL spikes. The provider is degrading mid-stream |
| Cache hit rate | Prefix and response cache effectiveness | Drops sharply. Cost and latency regress |
| Failover rate | How often requests leave the preferred route | Rises outside a planned rollout. Check health, policy and quality |
| Shed rate | How much load you reject | Rises above the expected baseline or stays elevated |
| Per-provider error rate | Which upstream returns failures | One provider moves above its error baseline |
Failover rate and shed rate can reveal routing or capacity pressure before overall availability moves.
I would set separate availability, p99 latency and cost targets for each product area. I would manage them with an error budget and pause risky routing changes when it is exhausted. Failover rate shows traffic leaving the primary provider. Shed rate shows active rejection. Both can move before a coarse availability number changes.