2 min lesson
Evaluation and A/B testing at scale
Consider this situation: "A new completion model raises acceptance rate by 8% in an A/B test. Why might you still not ship it and what would you check?" Start with the decision, then the evidence.
Step 1 of 3
Cursor runs experiments on millions of users to advance the agent. That sentence is in the job description and it means you'll be asked: how do you actually know one model or prompt is better than another?
The trap is measuring what's easy instead of what matters. Acceptance rate is trivial to log and dangerously misleading on its own. The skill is building a metric stack that catches the gap between "the user clicked accept" and "the code was good."