2 min lesson
Run relevant checks inside the loop
Start at the first move in "Run relevant checks inside the loop" and carry it through to the proof.
Step 1 of 2
Run relevant checks inside the loop
Code supports useful mechanical checks, but each check has limits. A successful build proves that the code compiled in that environment. It does not prove the behavior is correct, secure or complete.
- 1Run parsing, type checking and the build steps that cover the changed language and package.
- 2Run focused tests for the changed behavior, then broader tests when the risk and available time justify them.
- 3Feed check failures back into the loop with their command, exit status and relevant output.
- 4State which behavior, integration or security assumptions remain unverified and place them beside the diff for review.
Show the exact diff, the checks that ran and the checks that did not. Keep rejection and checkpoint recovery available. A user should not need to reconstruct the agent's hidden state to understand or undo a weak change.
Name the failure first, then match it to a control. Missing API context needs retrieval. A malformed edit needs schema and anchor validation. A type error needs the type checker. A behavior error needs a relevant test or review. State what remains after those checks.
I would read the current symbols and tests before editing, validate the edit format and anchors, then run the relevant type checks and tests. I would feed actual failures back into the loop. The final diff would show what ran and what remains unverified so the user can review or reject it.