Gemini 3.8 Flash: Google’s Bargain Price Against Claude, and the Catch

Reihen von Server-Racks in einem Rechenzentrum mit blauer Beleuchtung
Photo by Marc PEZIN on Unsplash

Google released Gemini 3.8 Flash on September 2, its third budget model in six weeks. The introductory prices undercut Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.6 Sol by a wide margin, and Google is adding a dedicated cybersecurity edition. Competition among AI models is shifting further away from raw peak performance and toward cost per task. Read the numbers closely, though, and the low token price does not automatically translate into a lower bill.

Key takeaways

  • During the introductory phase, Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens; both figures double on January 1, 2027.
  • For comparison, Claude Opus 5 runs at $5 and $25, and GPT-5.6 Sol at $4 and $20 per million tokens.
  • Despite the same token price as its predecessor, the cost per task rose about 40 percent according to analysis firm Artificial Analysis, because the model consumes more reasoning steps and tool calls.
  • A second variant called Gemini 3.8 Flash Cyber is trained to find and fix software vulnerabilities and is available only through a partner program.
  • Google’s actual frontier models, Gemini 3.5 Pro and Gemini 4, still have not shipped.

One model, two editions

The standard version of Gemini 3.8 Flash targets programming, AI agents, and multi-step reasoning. On the DeepSWE v1.1 benchmark, which measures longer-running software tasks, the model scores 73.7 percent, just behind Claude Opus 5 at 74.0 percent and ahead of GPT-5.6 Sol at 72.7 percent. The predecessor, Gemini 3.7 Flash, scored 65.3 percent here. On the Intelligence Index from Artificial Analysis, a composite of several tests, Gemini 3.8 Flash rises from 56 to 59 points. The official model card from Google DeepMind lists a context window of one million input tokens and a maximum of 64,000 output tokens, a knowledge cutoff of March 2026, and known weaknesses: occasional hallucinations and sporadic timeouts under heavy load.

The second edition, Gemini 3.8 Flash Cyber, specializes in vulnerability analysis. It scores 86.2 percent on the CyberGym benchmark, ahead of GPT-5.5-Cyber at 85.6 percent; on automated patching in CWE-Bench it reaches 47.2 percent, just behind Anthropic’s Fable 5. Against prompt injection, meaning hidden instructions inside input data, Google reports an attack success rate of 5.5 percent. How quickly such agent systems can slip their leash was shown recently by Anthropic’s experiment with a deliberately manipulative model.

The price that isn’t one

Gemini 3.8 Flash costs exactly as much as its predecessor 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. These introductory prices apply through the end of 2026, after which they rise to $1.50 and $7.50. Even the higher figure stays well below Claude Opus 5 and GPT-5.6 Sol. So much for the headline. But Artificial Analysis also measured the actual cost of a standard task, and it rose from $0.40 on the 3.7 Flash to $0.58 on the new model, an increase of about 40 percent. The reason lies in how it works: Gemini 3.8 Flash tries harder, adding extra reasoning steps and calling tools repeatedly. Each of those steps consumes tokens, and in aggregate the added consumption partly cancels out the lower unit price. For applications where every cent matters, the predecessor can remain the cheaper choice.

Flash Cyber and the Fairwind program

Google does not release the Cyber variant to everyone. Access runs exclusively through the newly launched Fairwind program, aimed at government agencies, operators of critical infrastructure, and the maintainer teams of large software projects. According to Google, more than 650 partners are already involved, among them the security firm CrowdStrike and the data platform provider Snowflake. Inside the program the safety filters run less strictly, so partners can use the model for defensive work such as penetration testing. Anthropic and OpenAI operate comparable controlled channels. The approach is an admission: a model that reliably finds gaps can also exploit them, so the provider itself decides who gets it at full strength.

What this means for users

Google itself frames the Flash series as a price-performance offering, while its actual flagship models Gemini 3.5 Pro and Gemini 4 keep users waiting. New DeepMind head Koray Kavukcuoglu stresses that the company pursues both goals at once, yet for now Google is delivering speed in the budget segment and no new flagship. For developers in German-speaking countries this means capable models at low list prices are now plentiful, but the bill depends on actual token consumption, not on the price tag. Anyone who values data control should also check the open competition: Ollama Cloud runs current open-source models such as Qwen 3.5, DeepSeek V4, GLM-5.3, or Kimi K3 as an alternative to the major providers, in some cases at comparable cost. And as with the leaked GPT-6 Astra benchmarks, a single test score says little about everyday use. What matters is how a model solves the specific task, how often it double-checks along the way, and what shows up on the invoice at the end of the month.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top