2 min lesson
The structure that works for all four
Start at the first move in "The structure that works for all four" and carry it through to the proof.
Step 1 of 3
The structure that works for all four
- 1Restate and scope. Pin down what “good” means and what you're explicitly not solving. Ambiguity is the test; naming your assumptions is the answer.
- 2Propose the design. Sketch the concrete pipeline - data in, signal out - at a level you could start building.
- 3Attack your own design. Name the failure mode (reward hacking, contamination, off-policy bias) before they do and the mitigation for each.
- 4Define success. State the metric and the validation that would tell you it worked, including how you'd separate real gains from noise.
Learn more
Full explanation
Worked sketch - Drill 2, online RL for a Tab-style feature
Worked sketch - Drill 2, online RL for a Tab-style feature
- Data collection
- Log the context, suggestion and the user's accept, reject or edit signal without exposing a bad policy to all live traffic.
- Reward attribution
- Turn delayed behavior over the session into credit for the suggestion that caused it, instead of treating every accept as equal.
- Off-policy correction
- Use weighting or clipping because the logged sessions came from an older policy, not the model you are training now.
- Safety gates
- Use guarded exploration, held-out and adversarial checks, a KL leash and a rollback path before widening live exposure.
Lead every design answer with the failure mode, then the mitigation. The same move applies to Drill 1: for a non-verifiable grader, blend the learned signal with compile, lint and partial-test checks; score partial progress; hold adversarial reward-hacked stubs; and retrain on fresh policy samples with a KL leash. Anticipating proxy gaming unprompted is the clearest signal that you have trained these systems, not just read about them.
Generic RL design reads like a textbook. Tie the online-RL drill to the published Tab approach and the grader drill to non-verifiable coding rewards and reference the constraints they actually face - async multi-region pipelines, sandboxed environments, training on real sessions. Grounded specifics separate a candidate who studied Cursor from one who studied RL.