gemini-3.7-flash is live at Google’s official price — half of what the previous generation costs
Google’s next-gen Flash workhorse shipped 13 August, roughly three weeks after 3.6 Flash. Positioned as its “most intelligent workhorse model”, the release puts its weight on coding and agents: DeepSWE v1.1 rises from 48.6% to 65.3%, FrontierCode 1.1 from 34.4% to 43.6%, Terminal-bench 2.1 from 78.0% to 85.8%. On AutomationBench, a real business-workflow benchmark, it moves from 17.0% to 30.4% — ahead of Claude Sonnet 5 (10.7%) and GPT-5.6 Terra (23.6%) in Google’s own comparison. Long-context retrieval reaches 97.0% on GDM-MRCR v2, and PDF comprehension goes from 22.0% to 34.0%.
Specs and integration: 1M context / 64K output, multimodal input (text, image, audio, video), and three tunable thinking levels — low / medium (default) / high. Both the OpenAI-compatible and native Gemini endpoints work; migrating from 3.6 Flash is a model name change.
Pricing is identical to Google’s official rates (per 1M tokens):
- Input (prompt): $0.7500
- Output (completion, incl. thinking): $3.7500
- Cache read: $0.0750
- Cache write (5m): $0.7500
← Back to Live Updates · 📚 Monthly Archive