1 min lesson
CursorBench and reward shaping
Say what this means in practice: "CursorBench is the in-house answer to every limit those benchmarks have."
Step 1 of 3
The public benchmarks in the last section are a sanity floor, not the thing a Cursor release turns on. CursorBenchCursor's internal benchmark that scores models on both task performance and token efficiency, not accuracy alone. Press Enter for the full definition. is the in-house answer to every limit those benchmarks have - and the way it scores reflects the non-verifiable-reward problem this whole module started from.
CursorBenchCursor's internal benchmark that scores models on both task performance and token efficiency, not accuracy alone. Press Enter for the full definition. is Cursor's internal eval, built to be realistic and uncontaminated. Its tasks are drawn from real engineer queries - the actual requests Cursor's own engineers make of the agent - rather than scraped GitHub issues or leetcode prompts. Because the tasks are private and freshly sourced, a model can't have memorized the fix from pretraining, so a gain on CursorBench is closer to capability than recall.