Agentic search and document processing
Good first candidate when request volume, latency, and unit cost dominate.
Test retrieval quality, structured output validity, and retry rate.
Google model field guide · checked July 29, 2026
Gemini 3.5 Flash-Lite targets high-volume, latency-sensitive production traffic. The useful question is whether it meets your acceptance bar with fewer total resources—not whether its token price is the lowest number on a page.
Input at launch
$0.30 / MTok
Output at launch
$2.50 / MTok
Flowith status
Verify live
Good first candidate when request volume, latency, and unit cost dominate.
Test retrieval quality, structured output validity, and retry rate.
Use configurable thinking levels to trade speed for multi-step work.
Measure accepted patches and reviewer minutes, not code volume.
Google lists computer use as a built-in tool.
Validate permissions, target-site behavior, recovery, and human confirmation.
Low cost does not remove review requirements.
Use domain checks, refusal handling, logging, and escalation.
Evaluation
Build a representative set, then record p50 and p95 latency, output tokens, tool calls, retries, schema failures, accepted outputs, and reviewer minutes. Repeat the same set on 3.6 Flash before selecting a default.
Google says minimal and low thinking can prioritize speed and cost, while higher levels can support multi-step subagent work. Treat the level as part of the evaluated configuration and re-test after provider changes.
Gemini 3.5 Flash-Lite is Google's July 2026 high-throughput Flash model for low-latency and cost-sensitive agentic work.
Google published $0.30 per million input tokens and $2.50 per million output tokens. Verify the live API pricing page because provider pricing can change.
No. Provider rates and speed are inputs. Retries, longer outputs, tool loops, failures, and human review determine completed-task cost.
This page does not assert Flowith availability. Check the live workspace model selector for the current answer.