Google launches Gemini 3.6 Flash to cut token costs

Google released Gemini 3.6 Flash and 3.5 Flash-Lite to lower latency and token costs for enterprise AI agents. Google reports 3.6 Flash generates 17% fewer output tokens than 3.5 Flash.

Google introduced two new models this week, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, aimed at reducing token costs and latency for enterprise AI agents in production workflows. Google reports the 3.6 Flash model produces 17% fewer output tokens than the prior 3.5 Flash overall, with larger drops in some synthetic tests.

Gemini 3.6 Flash is positioned for coding and multimodal reasoning. Google published benchmark results showing a 49% success rate on the Datacurve DeepSWE test versus 37% for the earlier model, MLE Bench improving from 49.7% to 63.9%, and a GDPval-AA v2 score rising to 1,421 from 1,349. Google also reported token-use reductions of up to 65% on certain synthetic workloads. Pricing for 3.6 Flash is $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.

Gemini 3.5 Flash-Lite targets high-throughput, low-latency tasks such as document processing and background agent work. Google measured the model at 350 output tokens per second and reported a GDM-MRCR v2 long-context success rate of 72.2% compared with 60.1% for the prior 3.5 model. The model’s GDPval-AA v2 score rose to 1,140 from 642. Pricing for Flash-Lite is $0.30 per 1 million input tokens and $2.50 per 1 million output tokens.

Google added a native client-side computer-use tool to the Gemini API and Gemini Enterprise platforms to let models interact with an operating system without custom intermediary software. The company reported an OSWorld-Verified score of 83.0%, up from 78.4%, and said updated safeguards improved resistance to jailbreak attempts and lowered risks related to chemical, biological, radiological and nuclear misuse without raising refusal rates for benign requests.

A restricted variant, Gemini 3.5 Flash Cyber, is offered only to governments and vetted partners through a pilot for vulnerability remediation. Google reported competitive results on the CyberGym benchmark and described a deployment inside its CodeMender security agent in which multiple model instances cross-check findings before a human reviewer signs off on the final remediation.

Several enterprise customers have begun integrating the new models. Figma integrated 3.6 Flash into its prototyping infrastructure; Matt Colyer, Figma’s director of product engineering, called the model “a faster route through design iterations without a drop in output quality.” Legal tech firm Harvey and research tool Hebbia route multimodal document data through 3.6 Flash to ingest filings, parse document structure, read embedded charts and draft reports for human review.

Google said developers can access the new models through the Gemini API in Google AI Studio, Android Studio and the Gemini Enterprise Agent Platform. Consumers will see the models in the Gemini app, and Google indicated 3.5 Flash-Lite is rolling out in Search. Google also said Gemini 3.5 Pro remains in partner testing and that pre-training work for the next Gemini 4 architecture is under way.

Content on BlockPort is provided for informational purposes only and does not constitute financial guidance.
We strive to ensure the accuracy and relevance of the information we share, but we do not guarantee that all content is complete, error-free, or up to date. BlockPort disclaims any liability for losses, mistakes, or actions taken based on the material found on this site.
Always conduct your own research before making financial decisions and consider consulting with a licensed advisor.
For further details, please review our Terms of Use, Privacy Policy, and Disclaimer.

Articles by this author

This site is registered on wpml.org as a development site. Switch to a production site key to remove this banner.