Cursor Basics
Gemini 3.6 Flash: Previous Pricing and the 3.7 Upgrade
Gemini 3.6 Flash is the previous Google speed-tier model in Cursor. Cursor still lists it, hidden by default, at $1.50 per million input tokens, $0.15 cached input and $7.50 output. Start new high-throughput work with Gemini 3.7 Flash, which succeeds it at lower published rates, and keep 3.6 for controlled comparisons or older policy references.

On this page
- What is Gemini 3.6 Flash in Cursor?
- What strengths did Cursor publish for Gemini 3.6 Flash?
- How much context does Gemini 3.6 Flash actually take?
- How much does Gemini 3.6 Flash cost in Cursor?
- How much does the cache discount really save?
- How did Gemini 3.6 Flash compare with Gemini 3.1 Pro?
- When would I still pick Gemini 3.6 Flash?
What is Gemini 3.6 Flash in Cursor?
Gemini 3.6 Flash was Google's speed-tier step between Gemini 3 Flash and Gemini 3.1 ProGoogle's frontier model in Cursor for work that combines images and code, especially UI implementation and visual code analysis. Press Enter for the full definition. in Cursor. Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition. now succeeds it as the current speed-tier choice, while Cursor's pricing table keeps 3.6 listed but hidden by default. This page preserves the older model's documented behavior and rates for teams that still need a comparison or policy reference.
Cursor rated 3.6 Flash on three axes: fast speed, low cost and high intelligence. Those labels explain its original role, but they no longer make it the default Flash recommendation. Open the Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition. guide for a new evaluation.
Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition. cuts the published input and cache-read rates in half and lowers output from $7.50 to $3.50 per million tokens. Keep 3.6 pinned only when model identity is part of a controlled comparison or an existing approval.
- Model ID
- gemini-3.6-flash
- Provider
- Google. A third-party model, drawing from the Other ModelsCursor's pool of supported third-party models from providers such as Anthropic, Google and OpenAI, with separate plan access and usage from Cursor Models. Press Enter for the full definition. pool
- Context
- The model card lists a 200k context window with a 1M maximum
- Capabilities
- Agent and Thinking
- Cursor's tiering
- Speed: fast. Cost: low. Intelligence: high.
- Headline rates
- $1.50 per million input tokens, $7.50 per million output
Per the model card at cursor.com/docs/models/gemini-3-6-flash, checked 2026-08-14.
High is not the top of Cursor's own vocabulary here. Claude Opus 5's card reads frontier on that same axis, and Cursor tiers that one medium on speed and high on cost. The safe reading of a fast, low-cost, high-intelligence card, I think, is that this model will hold a multi-step task together, and that Cursor is not pointing at it for the hardest reasoning in the picker.
This exact topic is a hands-on Lesson: Separate product, model and cost — about 7 minutes, free to read.
Rather do it than read about it? Run 11 interactive Cursor walkthroughs in a simulated editor. Free, no account needed.
What strengths did Cursor publish for Gemini 3.6 Flash?
Cursor publishes three strengths for this model, and the three turn out to be one argument: keeping a fast model useful once a task stops being a single edit. The list below is Cursor's, paired with what each line changes while you work.
- Cursor's claim
- Reasoning-capable Flash model
- What it changes in practice
- Better at multi-step coding tasks than Gemini 3 Flash, while staying fast
- Cursor's claim
- 90% discount on cached input tokens ($0.15/1M)
- What it changes in practice
- Strong for repeated context, such as sending the same large codebase across a session
- Cursor's claim
- 1M token context window
- What it changes in practice
- Fits substantial portions of a repository in a single request
| Cursor's claim | What it changes in practice |
|---|---|
| Reasoning-capable Flash model | Better at multi-step coding tasks than Gemini 3 Flash, while staying fast |
| 90% discount on cached input tokens ($0.15/1M) | Strong for repeated context, such as sending the same large codebase across a session |
| 1M token context window | Fits substantial portions of a repository in a single request |
Strengths as published by Cursor, checked 2026-08-14.
Two of those three lines only pay off together.
A million-token ceiling is affordable to use because cached input reads at $0.15 per million, and the reasoning claim is what makes it worth putting multi-step work through the cheap tier at all. On one fresh pass through a large repository, you get the ceiling and none of the discount.
Cursor does not publish a limitations list here, the way it does for Claude Opus 5, so there is no first-party statement of where this model falls down. My first instinct was to read the tiering as a stand-in, and that is only half right: the tiering tells you where Cursor placed the model and nothing about how it behaves. Opus 5's card names a specific habit, over-elaborating in long sessions. Treat the gap as missing information and find the edges yourself, on work you can check.
How much context does Gemini 3.6 Flash actually take?
Cursor's own page states the context limit two ways, and the two are not the same claim. Both figures are below exactly as published, because a large-repository workflow planned around the wrong one is expensive to discover mid-session.
- Model card
- Context window: 200k. Max context: 1M.
- Strengths list
- "1M token context window. Fits substantial portions of a repository in a single request."
Both from cursor.com/docs/models/gemini-3-6-flash, checked 2026-08-14.
The consistent reading is that 200k is the standard window and 1M is the ceiling the model can be run at, the same shape as Claude Opus 5's card (300k, with a 1M maximum). Cursor's strengths bullet states 1M flatly, though, so the two lines can fairly be read as disagreeing.
If a workflow depends on getting a million tokens into one request, check the current model card and what your own plan offers in the picker before you commit to it.
The ceiling is the wrong number to design toward, whichever of the two is live on your plan. A Cursor field engineer's rule of thumb puts the first slip in quality at 50 to 60% of a filled window, which is why our CLI guide says to compact before that point rather than after it.
A big window can also hurt on its own terms. More code in view means more for the model to read past, and the file that matters gets easier to lose. The cache discount makes that trap cheaper to walk into, since keeping the whole repository loaded costs a tenth of what fresh input would. Keep the working set to what the task actually needs, and watch the context indicator the way you would on an expensive model.
How much does Gemini 3.6 Flash cost in Cursor?
Usage lands in the third-party Other ModelsCursor's pool of supported third-party models from providers such as Anthropic, Google and OpenAI, with separate plan access and usage from Cursor Models. Press Enter for the full definition. pool, which Cursor meters apart from the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool and which Claude Opus 5 and the rest of the third-party picker bill against too. The allowance included there runs $20 on Pro, $70 on Pro Plus and $400 on Ultra. Cursor StartAn India-only individual plan centered on Cursor Models, with cloud and mobile access but without Other Models, Auto, Bugbot, Automations or the Agent SDK. Press Enter for the full definition. includes only Cursor Models, so it does not include Gemini 3.6 Flash. All rates below are per million tokens.
- Model
- Gemini 3.6 Flash
- Input
- $1.5
- Cache write
- Not listed
- Cache read
- $0.15
- Output
- $7.5
| Model | Input | Cache write | Cache read | Output |
|---|---|---|---|---|
| Gemini 3.6 Flash | $1.5 | Not listed | $0.15 | $7.5 |
Published rates at cursor.com/docs/models/gemini-3-6-flash, checked 2026-08-14. Cursor's table shows a dash in the cache-write column rather than a figure.
The cache-read rate is the one to plan around. At $0.15 against $1.50 for fresh input, repeated context costs a tenth of what it cost the first time, and Cursor frames the discount around exactly that case: a codebase you keep re-sending rather than a scattering of one-off prompts.
One cell in that table is a dash rather than a number. Cursor has not published a cache-write rate for this model, so do not read it as zero when you model the cost.
The allowance is what actually runs out, and a low rate only changes how fast. At $1.50 input against Opus 5's $5, the same pool lasts roughly three times as long. On an individual plan, you add on-demand usage at the same rates or move up a tier. Teams first moves a member who empties Other ModelsCursor's pool of supported third-party models from providers such as Anthropic, Google and OpenAI, with separate plan access and usage from Cursor Models. Press Enter for the full definition. to the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool; continuing with Gemini 3.6 Flash uses on-demand billing when enabled. Auto remains a Router choice with mode-specific billing, not the fallback pool.
How much does the cache discount really save?
It saves up to 90% on the input side, only on the input side, and less than that once a team plan adds the Token Rate. On an individual plan cached input reads at $0.15 per million against $1.50 fresh, so a session that keeps re-sending the same repository converges toward a tenth of what its first pass cost.
On Teams and Enterprise the arithmetic moves, and not in the direction the 90% headline suggests. Naming this model on a request adds the Cursor Token Rate of $0.25 per million, and Cursor states that fee covers cached tokens as well as input and output. Add it to both sides and the spread narrows from ten to one down to $0.40 against $1.75, a little over four to one. Budget from the second pair if your account is on either plan.
Switching to Auto clears the fee in one of its three modes only. Cursor exempts Auto Cost and the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool, while Auto Balance and Auto Intelligence carry the rate on any request the router sends to a third-party model, and this is one. A team that moves to Auto for cost reasons and leaves everybody on Balance keeps paying the $0.25 whenever a request lands outside the Cursor Models pool.
The discount is also easy to lose by accident.
Switching models mid-session breaks the cache. Which is the part that catches people out: a thread that jumps to a Pro model for one hard step and comes back pays fresh input rates on the way in, and the tenth-of-the-price arithmetic quietly stops describing your session. Unrelated one-shot prompts miss the discount for a duller reason, since there is nothing there to re-read.
If a bill lands higher than the rate card implied, start at the usage dashboard, which shows usage and token breakdowns instead of one total. On an Enterprise account there is a second line to rule out, since Cursor prices regional data residency as a 10% uplift on Model pricing for the models it lists as eligible.
How did Gemini 3.6 Flash compare with Gemini 3.1 Pro?
Yes on all three published rates, though how much cheaper depends on which rate your work leans on. Cursor's pricing index lists Gemini 3.1 ProGoogle's frontier model in Cursor for work that combines images and code, especially UI implementation and visual code analysis. Press Enter for the full definition. at $2 input, $0.20 cache read and $12 output, against $1.50, $0.15 and $7.50 for Gemini 3.6 Flash.
Line those up and the input columns sit close together: fifty cents per million between them on fresh input, five cents on cached input. The output rates carry the real distance, $7.50 against $12. So the gap between these two is mostly a function of how much code they write, not how much they read, and a cached-heavy session narrows it to a few cents per million read. Decide on how much reasoning the work needs, then look at the output column, because that is where the difference shows up on the invoice.
When would I still pick Gemini 3.6 Flash?
Start new speed-tier work with Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition.. Keep 3.6 only when you need to reproduce an older result, validate a migration, or honor an approval that names the exact model. The older positioning below still explains what a 3.6 result was designed to do.
- Situation
- New high-volume coding work with bounded reasoning
- Reasonable pick
- Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition.
- Why
- Current successor with lower published input, cache-read and output rates
- Situation
- Reproduce a result created on 3.6
- Reasonable pick
- Gemini 3.6 Flash
- Why
- Hold the model identity steady for the comparison
- Situation
- Migrate a model-specific evaluation or allow-list
- Reasonable pick
- Compare 3.6 with 3.7
- Why
- Measure behavior before changing the approved model
- Situation
- Simple, high-throughput work with no reasoning demand
- Reasonable pick
- Gemini 3 Flash
- Why
- The cheaper Flash tier this model sits above on price
- Situation
- Hard reasoning where the answer has to be right first time
- Reasonable pick
- A Pro or frontier-tier model
- Why
- Cursor tiers this one fast and low-cost, not top of the picker
- Situation
- The same tradeoff decided per request instead of per session
- Reasonable pick
- Cursor Router, on Teams and Enterprise
- Why
- Routing sends simple work to price-efficient models and hard work to capable ones automatically
| Situation | Reasonable pick | Why |
|---|---|---|
| New high-volume coding work with bounded reasoning | Gemini 3.7 FlashGoogle's latest speed-tier model in Cursor for high-throughput coding, with a 1M maximum context window and a 90% cached-input discount. Press Enter for the full definition. | Current successor with lower published input, cache-read and output rates |
| Reproduce a result created on 3.6 | Gemini 3.6 Flash | Hold the model identity steady for the comparison |
| Migrate a model-specific evaluation or allow-list | Compare 3.6 with 3.7 | Measure behavior before changing the approved model |
| Simple, high-throughput work with no reasoning demand | Gemini 3 Flash | The cheaper Flash tier this model sits above on price |
| Hard reasoning where the answer has to be right first time | A Pro or frontier-tier model | Cursor tiers this one fast and low-cost, not top of the picker |
| The same tradeoff decided per request instead of per session | Cursor Router, on Teams and Enterprise | Routing sends simple work to price-efficient models and hard work to capable ones automatically |
Mapped from Cursor's published positioning and cost tiering.
Model choice moves cost and reasoning depth. Every model in the picker gets the same agent tools, this one included: file and folder search, reading and editing, shell commands, the browser, fetching rules.
Team size changes who makes this call more than it changes the answer. On a small team it stays per-person, and the correction usually arrives late, once somebody's allowance is gone for the month. On Teams or Enterprise the same pick collects the Token Rate, and Cursor Router becomes available, so the tradeoff can be settled per request instead of per session. An admin chooses which routing modes members can select from Auto, so that call lands on everyone at once.
Composer 2.5The current Composer release, better at long-running tasks and at judging when a job needs a light touch versus deep work. Press Enter for the full definition. at $0.50 input and $2.50 output is the other end of that argument, and it bills against the Cursor ModelsCursor's plan and usage pool for models it labels as Cursor Models, kept separate from the third-party Other Models pool. Press Enter for the full definition. pool, which Cursor describes as carrying significantly more included usage. Where the work is routine enough that reasoning depth is not the question, that is the cheaper pool to spend.
Frequently asked questions
What happened to Gemini 3.6 Flash in Cursor?
Gemini 3.7 Flash succeeded it as Google's latest speed-tier model in Cursor. The current pricing table still lists Gemini 3.6 Flash, hidden by default, at $1.50 input, $0.15 cache read and $7.50 output per million tokens. Keep it for controlled comparisons or older model-specific approvals rather than as the default for new work.
How much does Gemini 3.6 Flash cost in Cursor?
Per million tokens, Cursor lists $1.50 input, $0.15 cache read and $7.50 output, with no cache-write rate shown in its table. It draws from the third-party Other Models pool, which starts with Pro at $20 of monthly usage. Cursor Start does not include Gemini 3.6 Flash.
How does the Gemini 3.6 Flash cache discount work?
Cursor lists a 90% discount on cached input tokens, at $0.15 per million against the $1.50 standard input rate. It pays off on repeated context, such as re-sending the same large codebase across a session, and does nothing for a series of unrelated one-shot prompts. Switching models mid-session also breaks the cache, so a thread that hops models pays fresh input rates on the way back in.
Does the Gemini 3.6 Flash cache discount still apply on a Teams plan?
It applies to the model rates, but the effective spread is smaller. On Teams and Enterprise, selecting this model directly carries the Cursor Token Rate of $0.25 per million, and so does an Auto Balance or Auto Intelligence request that the router sends to it. Cursor states the fee covers cached tokens as well as input and output, so adding it to both sides puts cached input at $0.40 against $1.75 for fresh input, closer to four to one than ten to one. Auto Cost and the Cursor Models pool are the exempt paths.
Why keep the Gemini 3.6 Flash guide after Gemini 3.7 Flash?
The 3.6 guide preserves the exact rates, context wording and cache behavior needed to reproduce older results, review a model-specific allow-list, or measure a migration. For new speed-tier work, start with Gemini 3.7 Flash and compare it with 3.1 Pro only when visual coding or frontier-tier reasoning matters.
Does Gemini 3.6 Flash have a 200k or a 1M context window?
Cursor's page states both. The model card lists a 200k context window with a 1M maximum, while the strengths list says a 1M token context window that fits substantial portions of a repository in a single request. The consistent reading is 200k standard and 1M maximum, matching the shape of other model cards, but confirm the current card before designing a workflow around a million-token request.
Should I start a new project on Gemini 3.6 Flash?
Start a new speed-tier evaluation with Gemini 3.7 Flash. Use 3.6 only when reproducing an older result or validating a migration. For hard long-horizon work, compare the current Flash model with a frontier choice such as Claude Opus 5 on the same bounded task.
Sources & last verified
- Cursor - Gemini 3.6 Flash model documentation
- Cursor - Gemini 3.7 Flash model documentation
- Cursor - Models and pricing index
- Cursor - Team pricing
Cursor ships frequently. Facts verified against primary sources on August 14, 2026.