Teams
Cursor Token Rate: The Fee Explained (and How to Avoid It)
The Cursor Token Rate is a flat $0.25 per million tokens that Teams and Enterprise accounts pay on third-party model requests. Auto Cost and the Cursor Models pool are exempt; that pool currently contains Grok 4.6, Grok 4.5, and Composer 2.5. Auto Balance and Auto Intelligence incur the fee only when the router lands on a third-party model. It counts input, output and cached tokens.

On this page
What is the Token Rate on my Cursor bill?
It is a flat infrastructure fee: $0.25 per million tokens on third-party model requests for Teams and Enterprise customers. Cursor describes it as covering search, custom model execution and infrastructure costs: the machinery around the model call, not the model tokens themselves, which bill at their own rates.
That coverage list and the thing that triggers the charge describe different work. Cursor's team pricing page itemises the fee as semantic search, custom model execution for Tab and Apply, plus infrastructure and processing costs, and nothing on that list is specific to a third-party model. Why the same infrastructure comes free on a ComposerCursor's own fast coding model, tuned for the editor and priced well below frontier models; the recommended day-to-day model for executing a plan. Press Enter for the full definition. request is not something the page explains, and I would not guess at it. What the list does settle is that the fee is priced on your token traffic, at the same $0.25 whether that traffic went to the cheapest model in the table or the most expensive.
- Request
- Auto, Cost mode
- Token Rate applies?
- No. Auto Cost is exempt whichever model it routes to
- Request
- Auto, Balance or Intelligence mode
- Token Rate applies?
- Yes, once the router lands on a third-party model
- Request
- Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. selected explicitly
- Token Rate applies?
- No. Models in the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool are exempt
- Request
- Grok 4.6A Cursor and SpaceXAI frontier model for complex coding and knowledge work; it is a selectable model and is separate from the Grok Bot product. Press Enter for the full definition. or Grok 4.5 selected explicitly
- Token Rate applies?
- No. Models in the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool are exempt
- Request
- Claude / GPT / Gemini selected explicitly
- Token Rate applies?
- Yes. A third-party model you named yourself
- Request
- Third-party model via BYOK
- Token Rate applies?
- Yes. Your own key does not exempt it
| Request | Token Rate applies? |
|---|---|
| Auto, Cost mode | No. Auto Cost is exempt whichever model it routes to |
| Auto, Balance or Intelligence mode | Yes, once the router lands on a third-party model |
| Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. selected explicitly | No. Models in the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool are exempt |
| Grok 4.6A Cursor and SpaceXAI frontier model for complex coding and knowledge work; it is a selectable model and is separate from the Grok Bot product. Press Enter for the full definition. or Grok 4.5 selected explicitly | No. Models in the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool are exempt |
| Claude / GPT / Gemini selected explicitly | Yes. A third-party model you named yourself |
| Third-party model via BYOK | Yes. Your own key does not exempt it |
Per cursor.com/docs/models-and-pricing, checked 2026-08-14.
The BYOK row is the one that catches finance teams. Moving the model bill to your own provider account is the move internal cost comparisons often treat as the end of the platform charge. It does not clear this one.
Exempting Auto Cost and the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool makes the fee a nudge away from pinning: a team that names a third-party model on every request pays it, while a team on Auto Cost or a Cursor Model does not.
Cursor does not state that as the intent, so treat it as an observation about the pricing. Auto Cost, Grok 4.6A Cursor and SpaceXAI frontier model for complex coding and knowledge work; it is a selectable model and is separate from the Grok Bot product. Press Enter for the full definition., Grok 4.5, and Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. are the current fee-free paths.
This exact topic is a hands-on Lesson: Separate product, model and cost — about 7 minutes, free to read.
Rather do it than read about it? Run 11 interactive Cursor walkthroughs in a simulated editor. Free, no account needed.
Do Auto requests avoid the Token Rate?
Only Cost mode is always exempt. Cursor's models and pricing page states that Auto Cost and the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool are exempt. Auto Balance or Auto Intelligence incurs the Token Rate only when the request routes to a third-party model. Anything written before Cursor Router shipped on 2026-07-22, including an earlier version of this page, may treat Auto as a blanket exemption. Check both the mode and routed model instead.
The failure mode is a team that reads "move everyone to Auto" as the fix, changes the default, and finds the fee still on the invoice. Cursor Router is on by default for Teams plans, and Balance and Intelligence stay selectable unless an admin removes them, so "everybody is on Auto now" does not tell you which mode anyone is running. Balance and Intelligence also bill per request at the routed model's rate instead of Auto Cost's bundled rate, which is where it hurts, since the change meant to remove a fee can raise the model line while it does so.
The fee is 25 cents per million tokens. Auto Cost bills output at $6 per million and Claude Sonnet 5Anthropic's medium-tier frontier model in Cursor, with a 200k standard context window and up to 1M context for longer tasks. Press Enter for the full definition. bills $10, so on a pinned Sonnet 5 thread the model line is still what moves the invoice.
My own first read of this fee was that it was a routing problem, fixable by getting the team onto Auto. That is worth demoting to a side effect. Auto Cost does clear the fee, but the reason to send routine work there is the bundled model rate, and the fee is the smallest number in that decision. If you are setting a team default, set it on the model rate and let the exemption come along behind.
Which tokens does it count?
All of them: for each eligible request the charge covers input tokens, output tokens and cached tokens, and it applies across all three usage categories: included usage, on-demand usage and BYOK. Two consequences follow that surprise teams:
- Cache hits are not free for this fee. The rate is the same $0.25 whether a token was read from cache or sent fresh, so the discount you get on the model side never reaches this line.
- Context size multiplies it. More context per turn means more input tokens per turn, and an agent session is many turns. The fee grows with how much context you move, not with how hard the task was.
Put the flat rate next to the per-model cache-read prices and the cached-token line stops looking like a technicality. Cursor's pricing table lists cache reads at $0.02 per million on GPT-5.6 Luna, $0.075 on Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition., $0.20 on Claude Sonnet 5Anthropic's medium-tier frontier model in Cursor, with a 200k standard context window and up to 1M context for longer tasks. Press Enter for the full definition. and $1 on Claude Fable 5Anthropic's highest-priced generally available Cursor model for long, complex work, with a 30-day retention opt-in for eligible privacy and enterprise accounts. Press Enter for the full definition.. Against Luna, the fee is 12.5 times what the cached tokens themselves cost. Against Gemini 3.7 Flash, it is about 3.3 times the model's cached-input charge. Against Sonnet 5 it adds 125%, and against Fable 5 a quarter on top. Output tokens run the other way: $0.25 against Sonnet 5's $10 output rate adds 2.5%. So the fee's share of your bill climbs as your cache-hit ratio climbs. Long threads on a cheap model with a warm cache push it highest.
- One 200k-token request
- 5 cents of Token Rate (0.2M × $0.25/M), before a cent of model charge.
- A 40-turn session at that size
- About $2, on top of the model bill for the same 8M tokens.
- 25 engineers doing that every working day
- Roughly $1,000 a month of Token Rate alone, at 20 working days.
Illustrative arithmetic on the published rate, not a quote; verify current terms on Cursor's pricing docs.
Those rows are the ceiling of a heavy pattern rather than a forecast. Nobody sends 200k tokens on every turn, and the same 8M tokens cost far more on the model line than on this one. The useful thing the arithmetic settles, for me, is whether you are looking at a rounding error or a line item worth a meeting, and at team scale on context-heavy work it is the second one.
How do I reduce or avoid the Token Rate?
Cursor names two exempt paths, Auto Cost and the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool, and the steps below are about getting onto one of them. They run in this order for a reason: the first one tells you whether the other three are worth doing.
- 1Find out who is pinning models. The usage dashboard tracks Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. separately from Other ModelsCursor's pool of supported third-party models from providers such as Anthropic, Google and OpenAI, with separate plan access and usage from Cursor Models. Press Enter for the full definition., and per-user spend sits in the admin dashboard. A pinned model is often a leftover personal default, not a considered choice.
- 2Check which Auto modeCursor's automatic model router with Cost, Balance and Intelligence options that trade price against model capability. Press Enter for the full definition. your team is on. Only Cost mode is exempt, and Cursor Router is on by default for Teams plans, so a team-wide move to Auto changes nothing on this line if members are selecting Balance or Intelligence.
- 3Standardise suitable work on Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition.. Grok 4.6A Cursor and SpaceXAI frontier model for complex coding and knowledge work; it is a selectable model and is separate from the Grok Bot product. Press Enter for the full definition., Grok 4.5, and Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. are exempt. Keep named third-party models for the cases where a specific model demonstrably wins.
- 4Don't assume BYOK dodges it. Bringing your own key moves the model bill to your provider account but the Token Rate still applies to those requests.
Measuring first is not a formality. If the fee traces back to two people, you do not need a policy, and a policy is what most teams reach for. It also protects the comparison. Change the default and the model mix in the same week and nothing on the next invoice is attributable to either. The blunt version of step one, locking the model picker to Auto, takes the pinned model away from whoever had a reason to pin it, and I would leave that setting alone until the dashboard says more than a handful of people are involved.
On a small team none of this needs an admin setting. Scale the rows above down to ten seats and the Token Rate is a few hundred dollars a month at the heavy end, which two conversations and a shared screen will move. At a few hundred seats the levers change, though the order does not. Routing preferences let an admin remove up to two of the three Auto modes from the picker, and a team-wide spending limit caps the month while you sort out habits. Per-member limits are Enterprise only, and Enterprise pooled usage is the case that changes the arithmetic, because one person's pinned Fable 5 traffic draws down a budget everybody shares.
Two Cursor pages describe the trigger at different resolutions. The team pricing page words it as all non-Auto agent requests. The models and pricing page, checked 2026-08-14, is the one that names Auto Cost specifically and puts Auto Balance and Auto Intelligence inside the fee only when they route to a third-party model. Take the more specific page, and re-check both before a rate goes into a budget or a policy.
The same page adds a 10% uplift on Model pricing for regional data residency, and says nothing about whether that uplift reaches the Token Rate. If you are in a residency region, ask before you model it.
What should I check if the Token Rate is higher than expected?
Start with the usage dashboard, which tracks included usage separately for Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. and Other ModelsCursor's pool of supported third-party models from providers such as Anthropic, Google and OpenAI, with separate plan access and usage from Cursor Models. Press Enter for the full definition.. The Token Rate rides on eligible third-party requests in the second pool, so if that is the one draining, a lot of tokens are going to third-party models.
From there it is per-user spend in the admin dashboard, or the spending data endpoint on the admin API if you would rather pull it into a spreadsheet, which is probably the faster route once more than a handful of people are involved. Since the Teams allowance is per seat and does not transfer between members, the total tells you less than the breakdown does, and a single pinned model shows up in the breakdown.
Then check the mode. Not the picker default, the mode. Cursor Router is on by default for Teams plans and off by default on Enterprise, and a team that turned it on without setting routing preferences left Balance and Intelligence selectable.
Spending limits are the backstop while you work through that, team-wide on Teams and per member on Enterprise.
Frequently asked questions
What is the Token Rate charge on my Cursor invoice?
A flat $0.25 per million tokens that Teams and Enterprise accounts pay on third-party model requests. Cursor lists it as covering semantic search, custom model execution for Tab and Apply, and infrastructure and processing costs, and it counts input, output and cached tokens across included usage, on-demand usage and BYOK.
Does the Token Rate apply to Auto requests?
Auto Cost is exempt. Auto Balance and Auto Intelligence incur the Token Rate only when the router sends a request to a third-party model. Grok 4.6, Grok 4.5, and Composer 2.5 in the Cursor Models pool are also exempt. Check both the mode and routed model before you count on the exemption.
Do I pay the Token Rate with my own API keys (BYOK)?
Yes. BYOK moves model token costs to your provider account, but the Token Rate still applies to third-party model requests made through Cursor.
Is the Token Rate or the model charge the bigger number?
The model charge, in almost every case. Cursor's pricing table puts Claude Sonnet 5 at $2 per million input tokens and $10 output against the flat $0.25 Token Rate. The exception is cached input: cache reads bill at $0.02 per million on GPT-5.6 Luna and $0.20 on Claude Sonnet 5, so on a thread that mostly reads from cache the fee can cost more than the cached tokens do.
Do individual Pro plans pay the Token Rate?
Cursor documents the Token Rate on Teams and Enterprise plans. The individual usage pools are described as third-party models charged at the model's API price, with no Token Rate named. Check your plan's usage page rather than assuming either way.
Are cached tokens really charged?
For this fee, yes. The charge explicitly encompasses input, output and cached tokens. Caching still saves you on the model-side rates; it just doesn't reduce the Token Rate.
Sources & last verified
Cursor ships frequently. Facts verified against primary sources on August 14, 2026.