Google model field guide · checked August 14, 2026

Gemini 3.7 Flash is Google's new agentic workhorse—not an automatic migration.

Google released Gemini 3.7 Flash as a GA model for coding and agents on August 13, 2026. Its stable ID, one-million-token context window and time-bounded introductory price make it a strong pilot candidate; production teams should still compare it with 3.6 Flash and Flash-Lite on accepted-task cost.

Provider record

The facts to verify before implementation

Stable model ID

gemini-3.7-flash

Google Cloud lists the model as generally available from August 13, 2026.

Context and output

1,048,576 / 65,536 tokens

The first number is the context window; the second is the documented maximum output.

Modalities

Text out; text, image, audio and video in

File, duration and region limits still depend on the live Google surface.

Flowith access

Verify in the live workspace

A Google launch does not establish availability in Flowith's current model selector.

Pricing window

Do not turn the introductory price into a permanent claim

Through December 31, 2026

$0.75 input · $3.75 output

Google's introductory price per one million tokens for the published model table.

Starting January 1, 2027

$1.50 input · $7.50 output

Google's listed standard price per one million tokens. Priority, Flex, batch and caching terms can differ.

Budget from the live Google pricing table and your actual configuration. Provider rates do not include failed workflow attempts, retries, tool charges or human review.

Choose by workload and lifecycle—not by launch order

Coding and multi-step agents

Pilot Gemini 3.7 Flash first.

Use representative repository tasks, tool calls and acceptance tests; treat Google's benchmark results as provider evidence, not a guarantee for your workload.

An existing 3.6 production workflow

Compare before migrating.

Run the same prompts and tools on both versions, then compare accepted outputs, retries, latency, tokens and reviewer effort. Google has not announced a 3.6 retirement date.

High-volume or latency-sensitive work

Include Flash-Lite in the trial.

Measure completed-task cost rather than token price alone, including schema failures, retries, tool loops and human correction.

Live API or model tuning

Choose another supported model or surface.

Google Cloud marks Gemini Live API and tuning as unsupported for this model record; recheck the capability table before implementation.

Lifecycle boundary

Google's August 14 lifecycle table places both 3.7 Flash and 3.6 Flash in its short-term availability group with no retirement date announced. That is a current status, not a promise of indefinite availability; keep the model ID configurable and recheck the table before a migration.

Capability and access boundaries

Supported

Structured output, function calling, code execution, context caching, URL context, Google Search and Google Maps grounding.

Preview boundary

Computer use is listed as a preview feature. Preview status is not a production-readiness or permission guarantee.

Thinking control

LOW, MEDIUM and HIGH are supported; MEDIUM is the documented default. MINIMAL returns an API validation error.

Not supported in this record

Gemini Live API, model tuning and fixed quota. Availability and consumption options vary by region and account.

At launch, Google points developers to the Gemini API through AI Studio, Android Studio and Antigravity; enterprises to Gemini Enterprise Agent Platform and the Gemini Enterprise app; and eligible individuals to Spark in supported countries. Each surface has its own plan, billing, region, quota and permission rules. None of those surfaces proves Flowith availability.

Continue the model decision

Official sources

  1. Google launch announcement — model role, dated release, introductory pricing and Google access surfaces.
  2. Google Cloud model record — stable ID, GA stage, limits, modalities, tools, regions and unsupported capabilities.
  3. Google Cloud pricing — introductory and standard price windows plus consumption-specific terms.
  4. Google model lifecycle table — current 3.7 and 3.6 retirement status.
  5. Google DeepMind model card — safety evaluations and provider limitations.

Gemini 3.7 Flash FAQ

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is Google's generally available August 2026 Flash model for coding and agent workflows. Google positions it as the primary agentic workhorse in the Gemini 3 family.

What is the Gemini 3.7 Flash model ID and context window?

The stable Google Cloud model ID is gemini-3.7-flash. Its documented context window is 1,048,576 tokens and its maximum output is 65,536 tokens.

How much does Gemini 3.7 Flash cost?

Google lists introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. It lists $1.50 input and $7.50 output starting January 1, 2027. Check the live pricing table for the applicable inference mode and terms.

Should I replace Gemini 3.6 Flash with 3.7 Flash?

Not without a matched test. Google reports improvements over 3.6 for coding and agent work, but it has not announced a 3.6 retirement date. Compare accepted output, retries, latency, tokens and review effort on your own workload.

Is Gemini 3.7 Flash available in Flowith?

This page does not assert Flowith availability. Verify the live Flowith workspace model selector before planning a workflow around this model.