1 min lesson
The offline–online gap
Use "A model can climb your offline eval and not move a single shipped metric" to say what you would do next.
Step 1 of 2
The offline–online gap
The deepest eval problem is that the benchmark and the product can disagree. A model can climb your offline eval and not move a single shipped metric - or get worse for users while the benchmark says it improved. Cursor's whole premise is that research is judged by production impact, not by the offline number.
A benchmark gain that doesn't move a shipped user metric is suspect until proven otherwise. The strongest eval practice maintains a chain: offline eval correlates with a known online metric (acceptance, retention, follow-up edits) and that correlation is itself monitored. When offline and online diverge, you trust the product and go fix the eval - because the product is the ground truth the eval is only approximating.