Details
- Google AI Developers announced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, positioning both models for production agentic workflows.
- Gemini 3.6 Flash is aimed at cutting token usage, reducing overall cost, and completing multi-step tasks with fewer reasoning steps and tool calls.
- Gemini 3.5 Flash-Lite is Google’s fastest and most cost-effective model for low-latency, high-throughput work such as repetitive tasks, parallel subagents, agentic search, and document processing.
- Both models are available now through the Gemini API in Google AI Studio and Android Studio, with 3.6 Flash also available in Antigravity and the Gemini Enterprise Agent Platform.
- Google’s product page says 3.6 Flash is built for frontier performance across agents and coding, while Flash-Lite is tuned for speed and throughput at scale.
- The launch follows Google’s broader push to make Flash models the default choice for cheaper, faster developer and enterprise automation.
Impact
The release strengthens Google’s position in the race to make agentic AI cheaper to operate, not just smarter. By emphasizing fewer tokens, fewer tool calls, and higher throughput, Google is targeting the operational bottlenecks that matter most in production deployments. That could pressure rivals such as OpenAI and Anthropic to sharpen their own price-performance story for high-volume workflows, especially in coding, document processing, and subagent orchestration.