SpaceRock

Gemini 3.7 Flash is out now: Google’s new coding workhorse

By Shaurya Sharma · · 1045 words

Topics: AI, AI Agents, AI Benchmarks, AI Coding, AI Models, AI Performance, AI Pricing, Artificial Intelligence

Gemini 3.7 Flash is out now: Google’s new coding workhorse

Artificial intelligence is evolving at an unprecedented rate. Google has recently announced a new model, the latest iteration of its Gemini series of AI models, called Gemini 3.7 Flash.

Google is promoting this release as a cost-effective, high-performance model that focuses on more specialised applications, such as coding, web development, and large-scale AI agent orchestration, rather than being a general-purpose assistant.

In short, this model is more effective at doing actual work.

If you find yourself needing to code something, Gemini 3.7 Flash’s enhanced coding capabilities enable you to create programs with significantly more complex instructions.

It could also help you debug and build interfaces much faster. Even if you are not an experienced programmer who can quickly write code, you can still benefit from Gemini 3.7 Flash’s abilities. You will not have to spend hours trying to understand why your code does not work or fix minor issues; you can simply ask Gemini 3.7 Flash to write the code for you and move on with your life.

Another significant improvement is that Gemini 3.7 Flash better handles AI agents than ever before.

An AI agent is essentially a chatbot that can perform moderately complex tasks on its own. Previously, such tasks had to be done by humans because regular chatbots could not handle them. However, an AI agent driven by Gemini 3.7 Flash can perform moderately complex tasks by itself, such as searching the web and utilising multiple tools to find and organise the information. Google claims that it “thinks more carefully through its responses”, which means that it can perform tasks that require more intricate planning and execution. In other words, Gemini 3.7 Flash will be able to get unstuck better when it encounters a roadblock and finds a new way around it.

You might have tried giving instructions to an AI before, only for it to ignore most of them. Chances are, Gemini 3.7 Flash will be less likely to make such mistakes. Following instructions is one of the critical competencies of Gemini 3.7 Flash, which was specifically designed to process exceptionally long instructions better and follow them to the letter

Last but not least, Gemini 3.7 Flash is part of Google’s “Flash” series of models, meaning that it offers higher performance at a lower price point.

In essence, it makes more financial sense to use Gemini 3.7 Flash for applications that require cutting-edge artificial intelligence, such as developing actual products, because it will be significantly cheaper than alternatives.

Benchmarks

On its own figures, Gemini 3.7 Flash is right around the mid-range frontier models in price for roughly double performance, and it is large enough coding gains over 3.6 Flash to be noted.

On coding and software engineering, in particular:
Gemini Benchmarks

FrontierCode 1.1 Main (production code quality): 43.6 per cent versus 34.4 per cent for 3.6 Flash, with Claude Sonnet 5 at 42 per cent and GPT-5.6 Terra with 41 per cent.

DeepSWE v1.1 (long-horizon software engineering): 65.3 per cent versus 49.0 per cent for 3.6 Flash, with GPT-5.6 Terra leading the pack at 69.6 per cent, Muse Spark 1.2 at 59.3 per cent, and Claude Sonnet 5 at 54.0 per cent.

Code Arena (web development, Elo): 1588 (first), with Claude Sonnet 5 at 1541, Muse Spark 1.2 at 1535, and GPT-5.6 Terra at 1523.

Terminal-bench 2.1: 85 per cent for Gemini 3.7 Flash, with GPT-5.6 Terra at 87 per cent (second). On the more difficult 3.0 version, everyone drops to 14.9 per cent (Gemini), 14.6 per cent (Claude Sonnet 5), and 20 per cent (GPT-5.6 Terra).

Enterprise workloads and agents:
SpaceRock

AutomationBench (business workflow automation): 30 per cent for Gemini 3.7 Flash versus 17.0 per cent for 3.6 Flash, 23.6 per cent for GPT-5.6 Terra, and 10.7 per cent for Claude Sonnet 5.

GDP.pdf (expert document comprehension): 34.0 per cent for Gemini 3.7 Flash versus 22.0 per cent for 3.6 Flash.

Harvey LAB-AA (complex legal workflows): 90 per cent (first in the pack).

OSWorld-2.0 (agentic computer use): 38per centent for Gemini 3.7 Flash, with GPT-5.6 Terra at 50.2 per cent leading the way.

Agent’s Last Exam: 26.3 per cent for Gemini 3.7 Flash versus 33.3 per cent for Claude Sonnet 5.

GDPVal-AA v2 (general knowledge work, Elo): 1525 for Gemini 3.7 Flash, with Muse Spark 1.2 leading the way at 1628 (first).

Long context and multimodal processing:

LVBench (long video understanding): 85.4 per cent (first).

GDM-MRCR v2 (long-context retrieval): 97 per cent at 128k tokens, and 62 per cent at 1M, with a 1,048,576-token input and 65,536-token output context window.

Composite and pricing:
SpaceRock

Artificial Analysis Intelligence Index: 56, which is better than Claude Sonnet 5 at 55 and its own predecessor at 52, but lower than GPT-5.6 Terra and Muse Spark 1.2 at 57.

Pricing: 0.75 dollars per million input tokens and 3.75 dollars per million output tokens for now; the introductory rate expires on 31 December 2026. From 1 January 2027, it is going to be 1.50 and 7.50 dollars per million input/output tokens, making it more expensive than Claude Sonnet 5 at 2.00 and 10.00, GPT-5.6 Terra at 2.00 and 12.00, and Muse Spark 1.2 at 1.25 and 4.25.

Conclusion

Overall, Gemini 3.7 Flash is a slight improvement with a higher price than its predecessors, which does not seem to be meant to compete at the very top. It scores first in coding and document processing, but the performance drops significantly on the open-ended agent tasks and general knowledge workloads like Agent’s Last Exam, GDPVal-AA v2, and OSWorld-2.0. More importantly, this model does not offer value in enterprise use cases over Claude, competing significantly with GPT-5.6 Terra and Muse Spark 1.2.

In particular, Gemini 3.7 Flash excels at structured programming and document processing but lags in general use cases. Terminal-bench 3.0 is more than brutal on all models, with the scores in the low teens for Gemini 3.7 Flash. The low price of 0.75 dollars per million input tokens may change as soon as 1 January 2027, when it is set to double (to 1.50 dollars per million input tokens). By shipping three iterations of the Gemini Flash model within a few months, Google is either trying to accelerate adoption or respond to competition, but the final judgment will be made when the next iteration of the Gemini Pro is available.

Add SpaceRock as a preferred source on Google