Pricing Index ·

AI model API pricing

Compare input, output, and cached-input rates across 38 AI model records tracked by benchr. Of those, 34 records have provider token pricing from OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, MiniMax, DeepSeek, Mistral, Meta, and other providers; unpriced media-model records remain available in the open data.

Data from models.json · verified August 21, 2026 No paid placement · commercial links are disclosed

Snapshot scope: On September 11, 2026 the lowest listed rates among hosted models tracked by benchr start at $0.15 input per 1M tokens: Qwen3.8-Flash lists $0.15 / $0.47, GLM-5.3-Flash $0.15 / $0.50 now that its launch promotion ended on September 9, and Mistral Small 4 $0.15 / $0.60. DeepSeek-V4.1-Flash replaced the retired V4-Flash on September 10 at $0.30 / $1.20 at peak, half that off-peak. This is not a market-wide guarantee; check the provider's live rate card before budgeting.

All models by price

Model Provider Input / 1M Output / 1M Cached input / 1M Details
Qwen3.6-27BAlibabaSelf-hostedSelf-hosted—Pricing detail
Llama 4 MaverickMetaSelf-hostedSelf-hosted—Pricing detail
Llama 4 ScoutMetaSelf-hostedSelf-hosted—Pricing detail
Phi-4MicrosoftSelf-hostedSelf-hosted—Pricing detail
Qwen3.8-FlashAlibaba (Qwen)$0.150$0.47—Model record
GPT-5 MiniOpenAI$0.250$2.00$0.025Full record
MiniMax M3MiniMax$0.300$1.20$0.060Model record
Gemini 3.5 Flash-LiteGoogle$0.300$2.50—Review
DeepSeek-V4.1-FlashDeepSeek$0.300$1.20$0.006Full record
Mistral Large 3Mistral$0.500$1.50—Pricing detail
Gemini 3.7 FlashGoogle$0.750$3.75$0.075Full record
Gemini 3.8 FlashGoogle$0.750$3.75$0.075Full record
Kimi K2.6Moonshot AI$0.950$4.00$0.160Pricing detail
Claude Haiku 4.5Anthropic$1.00$5.00$0.100Full record
GPT-5OpenAI$1.25$10.00—Pricing detail
Grok 4.3xAI$1.25$2.50$0.200Pricing detail
DeepSeek V4-ProDeepSeek$1.32$3.96$0.044Full record
GLM-5.2Z.AI$1.40$4.40$0.260Review
GLM-5.3Z.AI$1.40$4.40$0.260Review
Gemini 3.6 FlashGoogle$0.750$3.75$0.075Pricing detail
Gemini 3.5 FlashGoogle$1.50$9.00$0.150Pricing detail
Mistral Medium 3.5Mistral$1.50$7.50—Pricing detail
Gemini 3.1 ProGoogle$2.00$12.00$0.200Full record
Qwen3.8-MaxAlibaba (Qwen)$2.00$6.00—Model record
Grok 4.5xAI$2.00$6.00$0.500Review
Claude Sonnet 5Anthropic$2.00$10.00$0.200Full record
Grok 4.6xAI$2.00$6.00$0.500Review
GPT-5.4OpenAI$2.50$15.00$0.250Pricing detail
Kimi K3Moonshot AI$3.00$15.00$0.300Review
Claude Sonnet 4.6Anthropic$3.00$15.00$0.300Full record
GPT-5.6 SolOpenAI$4.00$20.00$0.400Full guide →
Claude Opus 5Anthropic$5.00$25.00$0.500Full record
Claude Opus 4.8Anthropic$5.00$25.00$0.500Full record
Claude Opus 4.7Anthropic$5.00$25.00$0.500Full record
GPT-5.5OpenAI$5.00$30.00$0.500Full record
GPT-6 AstraOpenAI$10.00$50.00$1.00Full record
Claude Fable 5.1Anthropic$10.00$50.00$0.250Full record
Claude Fable 5Anthropic$10.00$50.00$1.00Full record

The DeepSeek rows carry the peak rate, which is what a request pays unless it is deliberately scheduled elsewhere. Outside 01:00–04:00 and 06:00–10:00 UTC on weekdays the same models bill at half: $0.15/$0.60 for DeepSeek-V4.1-Flash and $0.66/$1.98 for DeepSeek-V4-Pro. The retired V4-Flash's last listed rates stay on the V4-Flash page.

Individual model pricing guides

Detailed breakdown per model: token costs, context limits, caching structures, use-case recommendations, and cost scenarios.

Pricing articles and analysis

Deeper reading on cost strategy, provider comparisons, and how to reduce your API bill.

How to read these pricing tables

API providers bill in tokens — roughly 750 words per 1,000 tokens for English text. Prices are listed per million tokens because per-token rates are too small to read ($0.000001250 is clearer as $1.25/1M). Most providers charge separately for input (text you send) and output (text the model generates).

Output tokens typically cost 4–12× more than input tokens because they require a separate forward pass per token. This means the input-to-output ratio of your workload matters enormously. A summarization task (high input, low output) has a very different cost profile than a story-generation task (low input, high output).

Three cost strategies for 2026

Route by task complexity. Use cheap small models (GPT-5.6 Luna at $0.20 input per 1M, DeepSeek-V4.1-Flash at $0.30 peak or $0.15 off-peak, Gemini 3.5 Flash-Lite at $0.30) for classification, routing, and simple extraction. Reserve expensive flagships for tasks that require their capability. Most pipelines can route 70–80% of calls to cheaper models.

Maximize prompt caching. If your system prompt or document context repeats across calls, caching cuts that portion of your input cost by 90% (Anthropic, Google) or ~99% (DeepSeek). This single optimization often reduces monthly bills more than switching models entirely.

Use Batch API for offline workloads. OpenAI, Anthropic, and others offer 50% discounts for asynchronous batch jobs — dataset evaluation, classification runs, bulk translation. If your workload can tolerate a 24-hour return window, batch pricing cuts your bill in half.

Use the Cost Calculator to model your specific token mix, or the Cheapest API leaderboard to find the lowest-cost model that meets your quality bar.

Frequently asked questions

Which tracked hosted AI API has the lowest listed price?

DeepSeek repriced on August 16, 2026 and replaced V4-Flash with DeepSeek-V4.1-Flash on September 10. Among hosted models tracked by benchr on September 11, 2026, three models list $0.15 input per 1M tokens: Qwen3.8-Flash at $0.47 output, GLM-5.3-Flash at $0.50 after its half-price launch promotion ended on September 9, and Mistral Small 4 at $0.60. Among the larger providers, GPT-5.6 Luna lists $0.20 / $1.20, and DeepSeek-V4.1-Flash lists $0.30 / $1.20 at peak and half that off-peak. This is not a market-wide guarantee; check the provider's live rate card before budgeting. Open-weight models like Llama 4 Scout and Phi-4 have no per-token licensing fee but require self-hosting infrastructure.

What does input vs output pricing mean?

API providers charge separately for input tokens (text you send to the model) and output tokens (text the model generates). Output tokens typically cost 4–10× more per token than input. Most real applications have 3–8× more input tokens than output.

What is prompt caching and how much does it save?

Prompt caching stores your static system prompt or document context so that repeated calls reuse it at a steep discount. Anthropic offers 90% off, Google 90% off, DeepSeek ~99% off, and GPT-5 cached input costs $0.125/1M — 90% below its $1.25 standard input rate. For agents with large repeated system prompts, caching can reduce your actual bill by 60–80%.

Are the model pricing pages on benchr accurate?

Prices are sourced directly from official provider documentation and verified against model-figures.json, benchr's single source of truth. The verified date is shown on each page. Prices change without notice — re-verify before making budget decisions.