Enterprise
AI Coding Security Checklist
An AI coding security checklist should cover data flow, model access, repo access, secrets, tool permissions, logging, human review and incident response. The question is simple: what can the tool read, what can it change and who reviews the result?
On this page
- What controls matter for AI coding security review?
- How should the rollout work?
- What's new in Cursor recently?
- Which checklist items are real boundaries and which are best-effort?
- Does a control you tested locally still hold in a cloud agent run?
- Why does network hardening sometimes make things less secure here?
- What retention numbers will a reviewer ask you for?
What controls matter for AI coding security review?
AI coding security review is a governance question before it is a tooling one. Four controls decide whether security can sign off: who has access, what the policy allows, how data moves, and whether anyone is measuring adoption. The table assigns each an owner.
- Control
- Identity
- Owner
- IT
- Evidence
- SSOSingle Sign-On. One company login (usually via SAML or OIDC) instead of a separate password per tool. Press Enter for the full definition., SCIMSystem for Cross-domain Identity Management. A standard for automatically creating and removing user accounts when people join or leave. Press Enter for the full definition. or membership source is defined
- Control
- Policy
- Owner
- Engineering leadership
- Evidence
- Allowed repos, tools and review rules are documented
- Control
- Security
- Owner
- Security team
- Evidence
- Data flow, secrets boundary and audit path are reviewed
- Control
- Adoption
- Owner
- DevEx
- Evidence
- Pilot metrics and training path are live
| Control | Owner | Evidence |
|---|---|---|
| Identity | IT | SSOSingle Sign-On. One company login (usually via SAML or OIDC) instead of a separate password per tool. Press Enter for the full definition., SCIMSystem for Cross-domain Identity Management. A standard for automatically creating and removing user accounts when people join or leave. Press Enter for the full definition. or membership source is defined |
| Policy | Engineering leadership | Allowed repos, tools and review rules are documented |
| Security | Security team | Data flow, secrets boundary and audit path are reviewed |
| Adoption | DevEx | Pilot metrics and training path are live |
The owners column matters more than the controls themselves, which I realise is a slightly odd thing to say about a governance table. But a control with nobody's name against it is a control nobody checks, and it fails at exactly the moment someone has to answer for it. The usual mistake is assuming security owns all four. They don't. Identity is IT's system of record, policy is an engineering-leadership call about what the team is allowed to ship, and adoption belongs to whoever owns developer experience. Security owns the data-flow question and the audit path, and they will ask about both.
Data flow is the one that stalls approvals, so answer it before the meeting rather than during it. Write down what the tool can read, what it can change, what leaves your network and where any of it is retained. Most of that comes straight from the vendor's documentation and your own configuration. It's dull work, but it turns a vague objection into a specific one, and specific objections are the kind you can actually close.
Interactive diagram. Tab through its regions; each focused region shows its detail in the panel below.
This is covered hands-on in Teams and Enterprise Admin — 6 short modules, free to read.
How should the rollout work?
A rollout works when it starts narrow and earns its expansion. One team, one repo and a few real tasks in week one, with policy and training in place before anyone scales it wider.
- 1Week 1: pick one team, one repo and three realistic tasks.
- 2Week 2: write the workflow standard from the pilot.
- 3Week 3: train champions and add policy guardrails.
- 4Week 4: expand only after quality, cost and review load are visible.
Four weeks is a shape, not a rule. The dates matter much less than the order, and the reason to start with a single team is that those first two weeks are mostly you finding out what your standard should even say. You can't write that from a pilot too broad to watch properly. Pick the team that will tell you when something doesn't work, rather than the team most likely to hand you a good result.
Expansion is where this usually goes wrong. Not because anyone's careless, either — the pilot went well, someone senior noticed, and now there's pressure. That's a hard thing to say no to. Before you add teams, though, check that you can answer three things with numbers instead of impressions: did review load go up or down, what does a developer actually cost per month, did quality hold. If any of those is a shrug, another fortnight of the same pilot beats a rollout you have to walk back.
What's new in Cursor recently?
Cursor ships often, so here is the current state of the surfaces this page touches. Each row links to the source where you can confirm the detail.
- Surface
- Compile 2026
- What to know
- Cursor's June 16 event highlighted Origin, larger from-scratch model training and Cursor Mobile alongside the broader June release wave.
- Surface
- Origin
- What to know
- Cursor's Origin page says code is moving faster than existing infrastructure was built to handle. The public page is waitlist-first, so migration and security details still need confirmation.
- Surface
- Model and mobile
- What to know
- Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. is available now. Cursor says a larger model is training with SpaceX. Mobile-native details remain beta until Cursor publishes a product page.
- Surface
- Automations
- What to know
/automate, Slack emoji triggers, GitHub issue/comment/review/workflow triggers, computer use, PR defaults and memory cleanup.
- Surface
- Cloud AgentsAgents that run in a Cursor-managed virtual machine, check out the repo, do the work and open a pull request, then shut down, with no load on your laptop. Press Enter for the full definition.
- What to know
- Guided cloud environment setup, reusable snapshots,
.cursor/environment.json,/in-cloud,/babysitand local/cloud handoff.
- Surface
- Review
- What to know
- BugbotCursor's automated PR reviewer that posts inline findings and can push fix commits from isolated VMs. Press Enter for the full definition. averages about 90 seconds and finds 10% more bugs per review, and can run locally before push with
/review. Cursor hasn't published Bugbot's underlying model. Don't assert one; treat the figures as perishable.
- Surface
- Design and Canvas
- What to know
- Design ModeA way to point at an element in Cursor's built-in browser and change it directly, instead of describing it in words. Press Enter for the full definition. supports multi-select and voice queueing; canvases support Design Mode, context reports, Debug with Agent, full-screen sharing and prompt buttons.
- Surface
- SDK and run modes
- What to know
- SDK agents can use custom tools, auto-review, JSONL/custom stores, nested subagents and request IDs; Auto-review Run Mode routes tool calls through safer execution paths.
- Surface
- Enterprise and pricing
- What to know
- Organizations sit above teams, groups scope model/spend/agent permissions and Teams now has Standard/Premium seats with Auto + ComposerCursor's own fast coding model, tuned for the editor and priced well below frontier models; the recommended day-to-day model for executing a plan. Press Enter for the full definition. and third-party API pools.
| Surface | What to know |
|---|---|
| Compile 2026 | Cursor's June 16 event highlighted Origin, larger from-scratch model training and Cursor Mobile alongside the broader June release wave. |
| Origin | Cursor's Origin page says code is moving faster than existing infrastructure was built to handle. The public page is waitlist-first, so migration and security details still need confirmation. |
| Model and mobile | Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. is available now. Cursor says a larger model is training with SpaceX. Mobile-native details remain beta until Cursor publishes a product page. |
| Automations | /automate, Slack emoji triggers, GitHub issue/comment/review/workflow triggers, computer use, PR defaults and memory cleanup. |
| Cloud AgentsAgents that run in a Cursor-managed virtual machine, check out the repo, do the work and open a pull request, then shut down, with no load on your laptop. Press Enter for the full definition. | Guided cloud environment setup, reusable snapshots, .cursor/environment.json, /in-cloud, /babysit and local/cloud handoff. |
| Review | BugbotCursor's automated PR reviewer that posts inline findings and can push fix commits from isolated VMs. Press Enter for the full definition. averages about 90 seconds and finds 10% more bugs per review, and can run locally before push with /review. Cursor hasn't published Bugbot's underlying model. Don't assert one; treat the figures as perishable. |
| Design and Canvas | Design ModeA way to point at an element in Cursor's built-in browser and change it directly, instead of describing it in words. Press Enter for the full definition. supports multi-select and voice queueing; canvases support Design Mode, context reports, Debug with Agent, full-screen sharing and prompt buttons. |
| SDK and run modes | SDK agents can use custom tools, auto-review, JSONL/custom stores, nested subagents and request IDs; Auto-review Run Mode routes tool calls through safer execution paths. |
| Enterprise and pricing | Organizations sit above teams, groups scope model/spend/agent permissions and Teams now has Standard/Premium seats with Auto + ComposerCursor's own fast coding model, tuned for the editor and priced well below frontier models; the recommended day-to-day model for executing a plan. Press Enter for the full definition. and third-party API pools. |
As of July 9, 2026. See Sources below for links.
Which checklist items are real boundaries and which are best-effort?
Sort every item into one of those two buckets before you sign anything. Deterministic controls block an operation whatever the model proposes: approvals, hooks, sandboxing, file system permissions. Best-effort controls only make the bad path less likely, and that group includes the Run Mode allowlist, .cursorignore, and team rules, which the model attempts to follow rather than obeys.
The item worth reading twice is .cursorignore. It excludes files from semantic search, agent file reading and context selection, and it is not a security boundary: a user can still open an ignored file by hand, and terminal and MCPModel Context Protocol. A standard that lets an AI agent pull in context from outside the repo, like Jira tickets or internal docs. Press Enter for the full definition. tools cannot honor it. Cursor's guidance is to pair it with approvals and file permissions, or to encrypt the data, on anything you would have to report.
The cost of getting this wrong isn't a breach on day one. It is a checklist that passed, filed somewhere, describing protections that don't hold, and the review that discovers it happens after an incident, at which point the document is evidence about your process rather than your controls.
Auto-review is the middle case. It runs allowlisted calls, sandboxes shell commands where it can and routes the rest through a best-effort classifier, which is exactly why Cursor recommends pairing it with hooks instead of treating it as the control.
Does a control you tested locally still hold in a cloud agent run?
Not automatically, and this is the row most checklists are missing. Cloud agents run command-based hooks from .cursor/hooks.json in your repository, plus team hooks and enterprise-managed hooks on Enterprise plans. User-level hooks in ~/.cursor/hooks.json are not loaded, because the cloud VM has no access to your home directory configuration.
So a hook an engineer verified on their own laptop may simply not exist in a cloud run. Prompt-based hooks don't run there either; cloud agents execute command-based hooks only.
More is missing than just the user-level hooks. beforeMCPExecution and afterMCPExecution are deferred for cloud agents. sessionStart and sessionEnd do not apply, and neither do the Tab hooks, since Tab completion is an IDE feature. Cloud agents also sometimes begin in a read-only environment for early exploratory turns, where hooks don't run at all; they start once the agent has a writable environment.
I had all of that filed as a configuration detail. It is a scope question. The honest version of this checklist has a column for where each control applies (local, cloud, or both), because a control covering one of them is a control you will describe wrongly in a review, and the review is the artifact that outlives the configuration.
Why does network hardening sometimes make things less secure here?
Because people route around it. Cursor's guidance is to allowlist *.cursor.sh and to exclude Cursor domains from SSL inspection, and the stated reason is the behaviour that follows when you don't: users disable security to make the tool work. That is a worse trade than the one you were trying to make.
Cloud AgentsAgents that run in a Cursor-managed virtual machine, check out the repo, do the work and open a pull request, then shut down, with no load on your laptop. Press Enter for the full definition. have their own outbound policy, with three access modes: Allow all network access, Default + allowlist, and Allowlist only. Enterprise admins can lock the team-level setting so members cannot override it. Allowlist only is the strictest and probably the one to start from if you can absorb the friction of adding destinations as they come up. MCPModel Context Protocol. A standard that lets an AI agent pull in context from outside the repo, like Jira tickets or internal docs. Press Enter for the full definition. network policy is set per server either way.
What retention numbers will a reviewer ask you for?
Three, and all of them are published. Indexed codebases expire automatically after 6 weeks of inactivity. Cloud Agent snapshots expire after 90 days of inactivity, with each start or resume extending the expiry by another 90 days, and Enterprise admins can cap Cloud Agent retention at Indefinite or 90 days, with custom windows in early access. Deleting an individual account removes that user's data, including indexed codebases, within 30 days.
Anyway, put those in the document rather than in a follow-up email. This is a section the reviewer can verify without your help, which is a reason to quote the numbers exactly instead of approximating them.
Frequently asked questions
Who is this guide for?
Security reviewers and platform teams approving AI developer tools.
What should I do next?
Start with one real repo task, capture the prompt and review the result before scaling the workflow.
Sources & last verified
Cursor ships frequently. Facts verified against primary sources on July 9, 2026.