BTC $63,085 โ–ผ0.50% ETH $1,879 โ–ฒ0.28% SOL $75.22 โ–ฒ 0.50% XRP $1.02 โ–ฒ 0.90%
Artificial Intelligence

Gemini 3.7 Flash Cuts Prices in Half, Beats Sonnet 5 on Coding

Gemini 3.7 Flash Cuts Prices in Half, Beats Sonnet 5 on Coding
๐Ÿ“‘ Table of Contents

Google is shipping Flash models faster than most rivals ship flagships โ€” and cutting prices while it does it. On August 13, Google introduced Gemini 3.7 Flash, just three weeks after Gemini 3.6 Flash, calling it "the most intelligent workhorse model yet for coding and agents." The new model arrives with an introductory price of $0.75 per million input tokens and $3.75 per million output tokens โ€” half of Gemini 3.6 Flash's standard pricing โ€” locked in through December 31, 2026.

That combination of iteration speed and aggressive pricing is aimed squarely at the agentic coding workloads where Google has been chasing OpenAI and Anthropic. And on several of the most-watched coding benchmarks, 3.7 Flash now leads โ€” though the "beats Sonnet 5" headline needs a footnote, and Google's own numbers come with caveats.

Three Weeks Between Flashes

The three-week gap between 3.6 Flash and 3.7 Flash is unusual even by Google's standards, and it signals a deliberate strategy: keep the workhorse tier moving, and let price do the talking. The intro rate of $0.75/$3.75 per million tokens holds until the end of the year; on January 1, 2027, it doubles to $1.50/$7.50, per the Google Cloud pricing page and the DeepMind model card.

For agent workloads โ€” where a single task can fan out into dozens or hundreds of model calls โ€” the halved price compounds. A team running 3.7 Flash for autonomous coding, browser automation or document work gets the same token budget at roughly half the cost of the previous generation, with a known price step at the new year. That predictability matters for teams that bill agent usage to clients or budget it per seat.

The Coding Wins, Scoped Honestly

The headline numbers are strong. On FrontierCode 1.1 Main, Gemini 3.7 Flash scores 43.6%, up from 34.4% for 3.6 Flash, and narrowly ahead of Claude Sonnet 5 (42.7%) and GPT-5.6 Terra (41.3%). On DeepSWE v1.1 it jumps to 65.3% from 49.0%, though GPT-5.6 Terra still leads that benchmark at 69.6%. On the WebDev Arena (Arena.ai) Elo ladder, 3.7 Flash sits at 1588 โ€” above 3.6 Flash's 1538, Claude Sonnet 5's 1541 and GPT-5.6 Terra's 1523.

BenchmarkGemini 3.7 FlashGemini 3.6 FlashClaude Sonnet 5GPT-5.6 Terra
FrontierCode 1.1 Main43.6%34.4%42.7%41.3%
DeepSWE v1.165.3%49.0%โ€”69.6%
WebDev Arena Elo1588153815411523

The gains extend beyond coding. Google reports AutomationBench at 30.4% versus 17.0% for 3.6 Flash โ€” beating Claude Sonnet 5's 10.7% and GPT-5.6 Terra's 23.6% โ€” and GDP.pdf at 34.0% versus 22.0%. These measure real agent behavior: automating software workflows and navigating document-driven tasks, not just writing functions in isolation.

But "beats Sonnet 5" needs context. Claude Sonnet 5 still leads Agent's Last Exam, 33.3% to 26.3%, and GPT-5.6 Terra leads Terminal-bench 2.1, 87.4% to 85.8%. Artificial Analysis places Gemini 3.7 Flash at 56 on its Intelligence Index, up from 52 but still below the 61-class frontier occupied by GPT-5.6 Sol, Grok 4.6 and Claude Opus 5. So the honest summary is: 3.7 Flash wins specific coding and automation benchmarks, and loses others โ€” it is a specialist workhorse, not a frontier claim. All deltas above are Google-reported figures awaiting independent verification.

Spark, Google's 24/7 Agent, Runs on Flash

The most visible proof of the model's readiness is consumer-facing. Gemini Spark โ€” Google's 24/7 personal agent, available on Google AI Pro and Ultra in more than 160 countries โ€” switched to 3.7 Flash from launch day. Google says the upgrade improves tool use across Workspace, from consolidating files and drafting emails to updating status documents.

Shipping a new model into a live agent with real users is a stronger vote of confidence than any benchmark table. It also means the pricing experiment has a captive showcase: every Spark interaction is a live test of 3.7 Flash's latency, reliability and cost profile at scale.

A Leadership Shuffle Behind the Speed

The release lands amid a reshuffle at the top of Google DeepMind. Demis Hassabis has stepped into a chair role as Alphabet chief scientist, while Koray Kavukcuoglu now runs DeepMind. The timing has also fed speculation that Gemini 3.5 Pro is delayed, though those rumors remain unconfirmed.

In that context, 3.7 Flash reads as more than a routine update: it is evidence that the workhorse tier โ€” and the agent products built on it โ€” can keep moving even while the company's leadership settles and its flagship roadmap faces questions. Google's bet is that most real-world agent work doesn't need the 61-class frontier; it needs a model that is fast, cheap and good enough to run thousands of times a day. With 3.7 Flash, Google is selling exactly that โ€” at half price, until the end of the year.

Benchmark and pricing figures above are Google-reported and have not been independently audited.

Sources

J

Jai

Jai covers trending tech, AI developments, and the cultural impact of emerging technologies at Veritya Daily. When he's not tracking viral stories, he's probably doom-scrolling through AI research papers.