Skip to lesson
Exit
Deep Dive: Graders, Rewards & Evals1 / 2

1 min lesson

The offline–online gap

Use "A model can climb your offline eval and not move a single shipped metric" to say what you would do next.

Step 1 of 2

The offline–online gap

The deepest eval problem is that the benchmark and the product can disagree. A model can climb your offline eval and not move a single shipped metric - or get worse for users while the benchmark says it improved. Cursor's whole premise is that research is judged by production impact, not by the offline number.

Tie the eval back to the product

A benchmark gain that doesn't move a shipped user metric is suspect until proven otherwise. The strongest eval practice maintains a chain: offline eval correlates with a known online metric (acceptance, retention, follow-up edits) and that correlation is itself monitored. When offline and online diverge, you trust the product and go fix the eval - because the product is the ground truth the eval is only approximating.