Gemini 2.5 Flash-Lite API Cost Calculator

Google's Gemini 2.5 Flash-Lite: $0.1/M input · $0.4/M output · $0.01/M cached input. Price your own workload below.

Cached input also incurs $1.00/MTok/hour storage.

Presets:

What Gemini 2.5 Flash-Lite costs in practice

WorkloadInputOutputCost
Short chat message500250$0.0001
10-page document + summary7,000500$0.0009
Agent / coding session50,00010,000$0.0090
1M tokens in + 1M out1,000,0001,000,000$0.500

Gemini 2.5 Flash-Lite vs nearest rivals

Closest-priced alternatives on a mixed 10K-input / 2K-output request:

ModelInput $/MTokOutput $/MTokMixed request
Gemini 2.5 Flash-Lite$0.1$0.4$0.0018
GPT-5 nano$0.05$0.4$0.0013
GPT-4o mini$0.15$0.6$0.0027
Mistral Small 4$0.15$0.6$0.0027

About Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite is the floor of Google's lineup by design — described by Google as the fastest and most budget-friendly multimodal model of the Gemini 2.5 family. The notable part is what survives at that price: full multimodal input and the same million-token-class context window as its bigger siblings, where competing budget tiers often cut context first.

Its natural work is the very-high-frequency layer — classification, tagging, routing, short summaries — where speed and per-request cost matter more than depth, with escalation to Flash or Pro for the requests that need more.

Flash-Lite is Google's cheapest text rate card, and unusually its output multiple over input is small, which makes it forgiving of verbose responses — the opposite of most models' cost profile. Cached input is discounted to near nothing per token but still accrues the hourly storage fee, so for bursty or low-volume traffic it is often cheaper to skip caching entirely and pay full input rate on a tiny number. At this price point the practical comparison is not other Google tiers but the cheapest rivals from other providers; the table above lines them up on a mixed workload.

Frequently asked questions

How much does the Gemini 2.5 Flash-Lite API cost?

Gemini 2.5 Flash-Lite costs $0.1 per million input tokens and $0.4 per million output tokens, with cached input at $0.01 per million (90% off). Note: Cached input also incurs $1.00/MTok/hour storage.

How much does a typical chat message cost with Gemini 2.5 Flash-Lite?

A short chat message (about 500 input and 250 output tokens) costs roughly $0.0001 with Gemini 2.5 Flash-Lite. At 1,000 such requests a day, that is about $4.50 per month.

Is Gemini 2.5 Flash-Lite cheap compared to similar models?

Gemini 2.5 Flash-Lite has the #2 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.0018 per request, versus $0.0013 for GPT-5 nano and $0.0027 for GPT-4o mini.

What does Gemini 2.5 Flash-Lite cost per 1,000 tokens?

Per 1,000 tokens, Gemini 2.5 Flash-Lite costs $0.00010 for input and $0.00040 for output, or $0.00001 for cached input. Providers quote rates per million tokens ($0.1/M in, $0.4/M out here), so divide by a thousand for the per-1K figure older pricing pages used.

How large a prompt can Gemini 2.5 Flash-Lite take?

Gemini 2.5 Flash-Lite accepts up to 1024K tokens of input context and can generate up to 64K tokens of output per request. At its input rate, filling the entire 1024K-token window costs about $0.105 per request before any output.

Related

Compare all models at once ·Count tokens in your prompt

All Google model costs: Gemini 3.8 Flash · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.1 Pro Preview · Gemini 2.5 Pro · Gemini 2.5 Flash

Last updated 2026-09-03. Prices verified against Google's official pricing page; see the methodology.