About o4-mini
o4-mini was the efficient half of OpenAI's final standalone o-series generation: a compact reasoning model for mathematics, coding and visual reasoning at volume, sitting under o3 the way mini tiers sit under flagships. It accepts text and images, produces text, and spends internal reasoning tokens before answering — the o-series signature.
OpenAI lists it as succeeded by GPT-5 mini, which folded that reasoning capability into the general-purpose family. It stays available for systems that standardised on its behaviour during the o-series era, especially high-throughput reasoning pipelines where its cost profile was the original draw.
Like every reasoning model, o4-mini bills its internal deliberation as output tokens, so effort level and problem difficulty move the bill more than prompt size does. Its cached-input discount lands midway between generations — deeper than the omni models offered, shallower than the GPT-5 family provides. GPT-5 mini, its listed successor, undercuts it on list price while adding the much larger family context window, so the succession here is unusually clean: migrate when convenient, and keep o4-mini only where validated behaviour matters. Batch processing remains supported for asynchronous reasoning workloads that can wait for results.
Frequently asked questions
How much does the o4-mini API cost?
o4-mini costs $1.1 per million input tokens and $4.4 per million output tokens, with cached input at $0.275 per million (75% off). Note: OpenAI lists o4-mini as succeeded by GPT-5 mini.
How much does a typical chat message cost with o4-mini?
A short chat message (about 500 input and 250 output tokens) costs roughly $0.0016 with o4-mini. At 1,000 such requests a day, that is about $49.50 per month.
Is o4-mini cheap compared to similar models?
o4-mini has the #14 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.020 per request, versus $0.015 for Gemini 3.8 Flash and $0.020 for Claude Haiku 4.5.
What does o4-mini cost per 1,000 tokens?
Per 1,000 tokens, o4-mini costs $0.0011 for input and $0.0044 for output, or $0.00028 for cached input. Providers quote rates per million tokens ($1.1/M in, $4.4/M out here), so divide by a thousand for the per-1K figure older pricing pages used.
How large a prompt can o4-mini take?
o4-mini accepts up to 195K tokens of input context and can generate up to 98K tokens of output per request. At its input rate, filling the entire 195K-token window costs about $0.220 per request before any output.
Related
Compare all models at once ·Count tokens in your prompt
All OpenAI model costs: GPT-5.6 Sol · GPT-5.6 Terra · GPT-5.6 Luna · GPT-5.4 · GPT-5 · GPT-5 mini · GPT-5 nano · GPT-4o · GPT-4o mini · o3
Last updated 2026-09-03. Prices verified against OpenAI's official pricing page; see the methodology.