Skip to lesson
Exit
Capstone: Mock Loop & Self-Exam1 / 3

2 min lesson

The structure that works for all four

Start at the first move in "The structure that works for all four" and carry it through to the proof.

Step 1 of 3

The structure that works for all four

  1. 1Restate and scope. Pin down what “good” means and what you're explicitly not solving. Ambiguity is the test; naming your assumptions is the answer.
  2. 2Propose the design. Sketch the concrete pipeline - data in, signal out - at a level you could start building.
  3. 3Attack your own design. Name the failure mode (reward hacking, contamination, off-policy bias) before they do and the mitigation for each.
  4. 4Define success. State the metric and the validation that would tell you it worked, including how you'd separate real gains from noise.
Learn more

Full explanation

Worked sketch - Drill 2, online RL for a Tab-style feature

Worked sketch - Drill 2, online RL for a Tab-style feature

How to answer it
Data collection
Log the context, suggestion and the user's accept, reject or edit signal without exposing a bad policy to all live traffic.
Reward attribution
Turn delayed behavior over the session into credit for the suggestion that caused it, instead of treating every accept as equal.
Off-policy correction
Use weighting or clipping because the logged sessions came from an older policy, not the model you are training now.
Safety gates
Use guarded exploration, held-out and adversarial checks, a KL leash and a rollback path before widening live exposure.
Interview move

Lead every design answer with the failure mode, then the mitigation. The same move applies to Drill 1: for a non-verifiable grader, blend the learned signal with compile, lint and partial-test checks; score partial progress; hold adversarial reward-hacked stubs; and retrain on fresh policy samples with a KL leash. Anticipating proxy gaming unprompted is the clearest signal that you have trained these systems, not just read about them.

Don't design in a vacuum, ground it in their work

Generic RL design reads like a textbook. Tie the online-RL drill to the published Tab approach and the grader drill to non-verifiable coding rewards and reference the constraints they actually face - async multi-region pipelines, sandboxed environments, training on real sessions. Grounded specifics separate a candidate who studied Cursor from one who studied RL.