2 min lesson
PPO, GRPO and the modern stack
Look at both sides of "PPO, GRPO and the modern stack" and name the signal that decides which one fits.
Step 1 of 2
PPO was the workhorse of RLHF. GRPO (Group Relative Policy Optimization) is the one that made large-scale RL on reasoning and code cheap enough to do at scale - and at Cursor, cheaper RL is the whole game.
Expect the screen to ask you to contrast them with real numbers and real memory footprints, not vibes. Both descend from the policy gradient in section one; they differ in how they estimate the advantage and how many models they keep in memory.
Interactive diagram. Tab through its regions; each focused region shows its detail in the panel below.
Step through each dimension - the difference is where the baseline comes from and how much memory you pay for it.