About Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite is the floor of Google's lineup by design — described by Google as the fastest and most budget-friendly multimodal model of the Gemini 2.5 family. The notable part is what survives at that price: full multimodal input and the same million-token-class context window as its bigger siblings, where competing budget tiers often cut context first.
Its natural work is the very-high-frequency layer — classification, tagging, routing, short summaries — where speed and per-request cost matter more than depth, with escalation to Flash or Pro for the requests that need more.
Flash-Lite is Google's cheapest text rate card, and unusually its output multiple over input is small, which makes it forgiving of verbose responses — the opposite of most models' cost profile. Cached input is discounted to near nothing per token but still accrues the hourly storage fee, so for bursty or low-volume traffic it is often cheaper to skip caching entirely and pay full input rate on a tiny number. At this price point the practical comparison is not other Google tiers but the cheapest rivals from other providers; the table above lines them up on a mixed workload.
Frequently asked questions
How much does the Gemini 2.5 Flash-Lite API cost?
Gemini 2.5 Flash-Lite costs $0.1 per million input tokens and $0.4 per million output tokens, with cached input at $0.01 per million (90% off). Note: Cached input also incurs $1.00/MTok/hour storage.
How much does a typical chat message cost with Gemini 2.5 Flash-Lite?
A short chat message (about 500 input and 250 output tokens) costs roughly $0.0001 with Gemini 2.5 Flash-Lite. At 1,000 such requests a day, that is about $4.50 per month.
Is Gemini 2.5 Flash-Lite cheap compared to similar models?
Gemini 2.5 Flash-Lite has the #2 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.0018 per request, versus $0.0013 for GPT-5 nano and $0.0027 for GPT-4o mini.
What does Gemini 2.5 Flash-Lite cost per 1,000 tokens?
Per 1,000 tokens, Gemini 2.5 Flash-Lite costs $0.00010 for input and $0.00040 for output, or $0.00001 for cached input. Providers quote rates per million tokens ($0.1/M in, $0.4/M out here), so divide by a thousand for the per-1K figure older pricing pages used.
How large a prompt can Gemini 2.5 Flash-Lite take?
Gemini 2.5 Flash-Lite accepts up to 1024K tokens of input context and can generate up to 64K tokens of output per request. At its input rate, filling the entire 1024K-token window costs about $0.105 per request before any output.
Related
Compare all models at once ·Count tokens in your prompt
All Google model costs: Gemini 3.8 Flash · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.1 Pro Preview · Gemini 2.5 Pro · Gemini 2.5 Flash
Last updated 2026-09-03. Prices verified against Google's official pricing page; see the methodology.