Google shipped Gemini 3.7 Flash this week, three weeks after Gemini 3.6 Flash, and cut the model’s API pricing by half through the end of the year. The company attributes the fast turnaround to developer feedback and algorithmic tuning. The discount has a hard expiration date, and that detail matters more than the release itself.
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through December 31. On January 1, 2027, both figures double to $1.50 and $7.50, the rate Gemini 3.6 Flash has charged since it launched. Context caching follows the same arc, rising from $0.075 to $0.15 per million tokens. A team running 100 million input tokens and 20 million output tokens today pays $150. Run the identical workload in February and the bill is $300.
Every benchmark figure below comes from Google’s own comparison table, worth flagging before citing any of it. On FrontierCode 1.1, a production-code-quality test, 3.7 Flash scored 43.6% against 42.7% for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra. It also leads Code Arena’s web-development Elo score, 1588 versus 1541 for Claude and 1523 for Terra, AutomationBench’s enterprise workflow tasks at 30.4% versus 10.7% and 23.6%, and GDP.PDF’s document-comprehension test at 34.0% versus 28.0% and 24.7%.
The picture flips elsewhere. GPT-5.6 Terra, OpenAI’s model, stays ahead on DeepSWE’s long-horizon software engineering test, 69.6% to 65.3%, and on Terminal-bench 2.1, 87.4% to 85.8%. Google’s own materials list Terra ahead on Terminal-bench 3.0 and OSWorld-2.0 too. Claude Sonnet 5, Anthropic’s model, leads Agent’s Last Exam, a multimodal desktop-and-OS benchmark, 33.3% to 26.3%. Google is not claiming a clean sweep, and the numbers do not support one.
Google has not shipped Gemini 3.5 Pro despite targeting it for June, and the flagship remains in partner testing. AI Insiders has covered the DeepMind leadership reshuffle that accompanied that delay separately, but a fast Flash cadence is what Google has to show while the larger release stays pending.
3.7 Flash is live now through the Gemini API in Google AI Studio and Android Studio, Google’s Antigravity development environment, the Gemini Enterprise Agent Platform, and Spark, the personal agent Google offers Gemini AI Pro and Ultra subscribers inside Workspace.
VentureBeat first reported the release, citing Google’s own benchmark and pricing disclosures. The outlet noted that Google’s results show a model substantially more competitive on coding and agent workloads while sitting in a lower price tier, not one that displaces every higher-priced rival outright.
The number worth tracking isn’t the per-token price. It’s cost per completed task. Whatever a coding or automation workload costs to run on 3.7 Flash today, budget for that figure to double the moment the calendar turns to January 1, 2027, and re-run the comparison against Claude Sonnet 5 and GPT-5.6 Terra before locking in a 2027 spend on an introductory rate that will not last.
VentureBeat, reported by Carl Franzen, published August 13, 2026.