Gemini 3.8 Flash API Cost Calculator

Google's Gemini 3.8 Flash: $0.75/M input · $3.75/M output · $0.075/M cached input. Price your own workload below.

Introductory pricing to 2026-12-31; from 2027-01-01 it is $1.50 in / $7.50 out. Cached input also incurs $0.50/MTok/hour storage.

Presets:

What Gemini 3.8 Flash costs in practice

WorkloadInputOutputCost
Short chat message500250$0.0013
10-page document + summary7,000500$0.0071
Agent / coding session50,00010,000$0.075
1M tokens in + 1M out1,000,0001,000,000$4.50

Gemini 3.8 Flash vs nearest rivals

Closest-priced alternatives on a mixed 10K-input / 2K-output request:

ModelInput $/MTokOutput $/MTokMixed request
Gemini 3.8 Flash$0.75$3.75$0.015
o4-mini$1.1$4.4$0.020
Claude Haiku 4.5$1$5$0.020
DeepSeek V4 Pro$1.32$3.96$0.021

About Gemini 3.8 Flash

Gemini 3.8 Flash heads the Flash tier of Google's Gemini 3 line. Google describes it as its most intelligent Flash model and points it at long-horizon software engineering, autonomous agents and complex enterprise workflows, while keeping the speed and cost profile that defines the tier. It is natively multimodal — text, images, video, audio and PDF documents go in, text comes out — and Google lists caching, code execution, file search and computer use among its supported capabilities.

It is a stable release under a plain model code rather than a preview alias, and sits above Gemini 3.7 Flash and Gemini 3.6 Flash in Google's listing; all three remain available on the same rate card, so the choice between them is about capability rather than price.

Google launched this model on the same time-limited introductory rates as the two Flash generations before it, with a stated date after which input, output and cache rates all double — the callout above gives the figures and the deadline. Because the three Gemini 3 Flash models share one rate card there is no discount for staying on an older one; the cost levers within the tier are elsewhere: batch and flex processing at half of standard, and context caching at a small fraction of the input rate plus an hourly storage charge that only pays off at a steady request volume. Thinking tokens bill as output, and Google's Pro tier prices well above this model with a further step up on very long prompts, so agentic workloads that fit Flash are where the value sits.

Frequently asked questions

How much does the Gemini 3.8 Flash API cost?

Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, with cached input at $0.075 per million (90% off). Note: Introductory pricing to 2026-12-31; from 2027-01-01 it is $1.50 in / $7.50 out. Cached input also incurs $0.50/MTok/hour storage.

How much does a typical chat message cost with Gemini 3.8 Flash?

A short chat message (about 500 input and 250 output tokens) costs roughly $0.0013 with Gemini 3.8 Flash. At 1,000 such requests a day, that is about $39.38 per month.

Is Gemini 3.8 Flash cheap compared to similar models?

Gemini 3.8 Flash has the #10 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.015 per request and $0.020 for o4-mini.

What does Gemini 3.8 Flash cost per 1,000 tokens?

Per 1,000 tokens, Gemini 3.8 Flash costs $0.00075 for input and $0.0037 for output, or $0.00007 for cached input. Providers quote rates per million tokens ($0.75/M in, $3.75/M out here), so divide by a thousand for the per-1K figure older pricing pages used.

How large a prompt can Gemini 3.8 Flash take?

Gemini 3.8 Flash accepts up to 1024K tokens of input context and can generate up to 64K tokens of output per request. At its input rate, filling the entire 1024K-token window costs about $0.786 per request before any output.

Related

Compare all models at once ·Count tokens in your prompt

All Google model costs: Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.1 Pro Preview · Gemini 2.5 Pro · Gemini 2.5 Flash · Gemini 2.5 Flash-Lite

Last updated 2026-09-03. Prices verified against Google's official pricing page; see the methodology.