Short answer

Ministral 3 14B currently has the lowest verified first-party output price in this catalog at $0.20 per million output tokens. On OpenRouter, Ministral 3 14B is the lowest among the 50 mapped models at $0.20 per million output tokens. Your cheapest total depends on input volume, output length, caching, and channel fees.

Which LLM APIs have the lowest token prices?

The table ranks verified first-party APIs by output-token price. The example workload uses 1 million input tokens and 250,000 output tokens with no cache discount. This makes the calculation repeatable, but it does not make the models equivalent in capability.

ModelOfficial inputOfficial outputOpenRouter in / outExample cost
1. Ministral 3 14BMistral AI $0.20 $0.20 $0.20 / $0.20 $0.25
2. DeepSeek V4 FlashDeepSeek $0.14 $0.28 $0.14 / $0.28 $0.21
3. Gemini 2.5 Flash-LiteGoogle $0.10 $0.40 $0.10 / $0.40 $0.20
4. Mistral Small 4Mistral AI $0.15 $0.60 $0.15 / $0.60 $0.30
5. DeepSeek V4 ProDeepSeek $0.435 $0.87 $0.435 / $0.87 $0.6525
6. MiniMax M2.7MiniMax $0.30 $1.20 $0.25 / $1.00 $0.60
7. MiniMax M2.5MiniMax $0.30 $1.20 $0.15 / $0.90 $0.60
8. Gemini 3.1 Flash-LiteGoogle $0.25 $1.50 $0.25 / $1.50 $0.625
9. Mistral Large 3Mistral AI $0.50 $1.50 $0.50 / $1.50 $0.875
10. Qwen 3.7 PlusAlibaba Cloud $0.40 $1.60 $0.32 / $1.28 $0.80

Sources: each official price links from the full model table; OpenRouter prices come from its public model feed. Example cost excludes taxes, tools, storage, regional premiums, and OpenRouter credit purchase fees.

What does “cheapest” actually mean?

A single price column cannot describe every workload. Retrieval and document-processing products may spend most of their budget on input tokens. Coding agents and long-form generation can spend much more on output. A model with cheap input and expensive output may be economical for classification but costly for generation.

Start with your monthly input tokens, output tokens, cache hit rate, and request count. Then compare the total bill. Cached-input discounts can materially change the result for repeated system prompts or shared document context, while tiered long-context pricing can make a nominal list price too optimistic for very large prompts.

When should you pay more for an LLM API?

Reliability

Production fallbacks

A higher rate can be rational when it includes dependable capacity, predictable limits, and a direct support relationship.

Capability

Harder workloads

Complex reasoning, tool use, and code generation may need a stronger model even when a budget model wins on raw price.

Governance

Data requirements

Regional processing, retention controls, and contractual terms can matter more than a small token-price difference.

How should you compare official APIs with OpenRouter?

Use official direct pricing when you want a first-party account, provider-specific features, and a direct commercial relationship. Use OpenRouter when a unified API, cross-provider fallbacks, and consolidated usage are more valuable. OpenRouter says inference pricing is passed through, while credit purchases currently carry a 5.5% fee with a $0.80 minimum. Review the OpenRouter pricing page before budgeting.

The full comparison should keep inference price and platform fee separate. That is why LLM API Prices lets you switch channels and optionally include the percentage fee instead of quietly adding it to every displayed token rate.

Cheapest LLM API questions

What is the cheapest LLM API by token price?

Among the verified first-party prices in this catalog, Ministral 3 14B has the lowest output-token list price at $0.20 per million output tokens. The answer can change when your workload has a different input-to-output ratio.

Is OpenRouter always cheaper than the official API?

No. Many OpenRouter inference rates match the first-party list price, while some routes are lower or higher. OpenRouter also charges a credit purchase fee, so compare the full workload cost instead of assuming one channel always wins.

Why are output tokens usually more expensive?

Generating output requires the model to run repeatedly for each new token. Providers therefore commonly charge more for output than input, which makes response length an important part of cost planning.

Should I choose an API only because it is cheapest?

No. Price does not capture response quality, latency, tool support, context limits, availability, data policies, or regional requirements. Use the cheapest model that still meets the requirements of the actual task.

Does this ranking include free tiers?

No. This table compares published pay-as-you-go token prices. Free quotas, promotional credits, batch discounts, and temporary offers are excluded unless a page explicitly says otherwise.