LLM API Pricing Cheat Sheet 2026: Real Costs Across 25+ Models (Updated)
2026 LLM API pricing for every major model: DeepSeek V4, GPT-5, Claude 4, Mimo, Gemini, Qwen, GLM, Kimi, Minimax, Hunyuan. Input/output costs per 1M tokens, context windows, and real-world cost analysis.
LLM API Pricing Cheat Sheet 2026: Real Costs Across 25+ Models (Updated)
Last updated: July 27, 2026 — Prices reflect real-time rates across all major LLM providers accessible via TokenPAPA. Bookmark this page for the latest data.
Finding the right LLM API for your project means comparing pricing across a dozen providers — and the price differences are staggering. A single API call to Claude Opus 4 costs 50x more than the equivalent call to DeepSeek V4 Pro.
We maintain this cheat sheet with real, verified pricing for every major model available on TokenPAPA. All prices are per 1 million tokens (input / output).
Quick Look: Cheapest to Most Expensive
| Rank | Model | Input Price | Output Price | Best For |
|---|---|---|---|---|
| 1 | Mimo V2.5 | $0.08 | $0.24 | Budget batch processing |
| 2 | DeepSeek V4 Flash | $0.14 | $0.42 | High-volume chat & coding |
| 3 | Mimo V2.5 Pro | $0.12 | $0.36 | Budget with quality |
| 4 | GPT-5.4 Mini | $0.15 | $0.60 | Fast, affordable chat |
| 5 | Qwen 3.5 Flash | $0.20 | $0.80 | High-speed inference |
| 6 | Gemini 3 Flash | $0.25 | $1.00 | Multimodal on budget |
| 7 | GLM-5 | $0.30 | $1.00 | Chinese-optimized tasks |
| 8 | Kimi K2.6 | $0.50 | $2.00 | Long-context (128K+) |
| 9 | Minimax M3 | $0.80 | $2.40 | Creative generation |
| 10 | DeepSeek V4 Pro | $0.28 | $0.84 | Best flagship value |
| 11 | Hunyuan HY3 | $1.00 | $4.00 | Enterprise Chinese NLP |
| 12 | Gemini 3.5 Flash | $1.25 | $5.00 | Fast multimodal |
| 13 | Claude Sonnet 4 | $3.00 | $15.00 | Balanced quality/speed |
| 14 | GPT-5.4 | $10.00 | $30.00 | High-quality reasoning |
| 15 | Claude Opus 4 | $15.00 | $60.00 | Best-in-class output |
| 16 | GPT-5.5 | $15.00 | $60.00 | Frontier intelligence |
Complete Pricing Table
Budget Tier (Under $1/M input)
| Model | Provider | Input | Output | Context | Category |
|---|---|---|---|---|---|
| Mimo V2.5 | Xiaomi | $0.08 | $0.24 | 128K | General chat |
| Mimo V2.5 Pro | Xiaomi | $0.12 | $0.36 | 128K | General chat |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.42 | 128K | Chat, coding |
| GPT-5.4 Mini | OpenAI | $0.15 | $0.60 | 128K | Fast chat |
| Qwen 3.5 Flash | Alibaba | $0.20 | $0.80 | 128K | Inference |
| Gemini 3 Flash | $0.25 | $1.00 | 1M | Multimodal | |
| DeepSeek V4 Pro | DeepSeek | $0.28 | $0.84 | 128K | Flagship |
| GLM-5 | Zhipu AI | $0.30 | $1.00 | 128K | Chinese tasks |
| Qwen 3.5 Plus | Alibaba | $0.40 | $1.60 | 128K | Balanced |
| GLM-5.1 | Zhipu AI | $0.50 | $1.80 | 256K | Advanced reasoning |
| Kimi K2.6 | Moonshot | $0.50 | $2.00 | 128K | Long context |
| GLM-5.2 | Zhipu AI | $0.50 | $2.00 | 256K | Deep reasoning |
| Minimax M2.5 | MiniMax | $0.60 | $2.00 | 128K | Creative |
| Minimax M2.7 | MiniMax | $0.70 | $2.20 | 256K | Creative pro |
| Minimax M3 | MiniMax | $0.80 | $2.40 | 128K | Generation |
Mid Tier ($1–$5/M input)
| Model | Provider | Input | Output | Context | Category |
|---|---|---|---|---|---|
| Hunyuan HY3 Preview | Tencent | $1.00 | $4.00 | 256K | Enterprise |
| Gemini 3.5 Flash | $1.25 | $5.00 | 1M | Fast multimodal | |
| Gemini 3 Pro Image | $2.00 | $8.00 | 2M | Multimodal pro | |
| Claude Sonnet 4 | Anthropic | $3.00 | $15.00 | 200K | Balanced |
Premium Tier ($5+/M input)
| Model | Provider | Input | Output | Context | Category |
|---|---|---|---|---|---|
| GPT-5.4 | OpenAI | $10.00 | $30.00 | 1M | High-quality |
| Claude Opus 4 | Anthropic | $15.00 | $60.00 | 200K | Best quality |
| GPT-5.5 | OpenAI | $15.00 | $60.00 | 2M | Frontier |
| Gemini 3.1 Pro | $10.00 | $40.00 | 2M | Enterprise multimodal |
Real-World Cost Examples
Let's put these numbers in perspective. Here's what common tasks actually cost:
Task: Generate a 500-word blog post (~700 output tokens)
| Model | Cost |
|---|---|
| DeepSeek V4 Flash | $0.00029 |
| GPT-5.4 Mini | $0.00042 |
| Claude Sonnet 4 | $0.0105 |
| GPT-5.5 | $0.042 |
Task: Analyze a 10-page document (~15K input tokens)
| Model | Cost |
|---|---|
| DeepSeek V4 Pro | $0.0042 |
| Mimo V2.5 Pro | $0.0018 |
| Claude Opus 4 | $0.225 |
| GPT-5.4 | $0.15 |
Task: Daily batch processing (1M input + 300K output tokens)
| Model | Daily Cost | Monthly Cost |
|---|---|---|
| Mimo V2.5 | $0.15 | $4.65 |
| DeepSeek V4 Flash | $0.27 | $8.16 |
| Gemini 3.5 Flash | $2.75 | $83.25 |
| Claude Sonnet 4 | $7.50 | $225.00 |
Price Comparison by Model Tier
Chinese Models (Best Value)
| Model | Input | Output | vs DeepSeek V4 Flash |
|---|---|---|---|
| Mimo V2.5 | $0.08 | $0.24 | 43% cheaper |
| DeepSeek V4 Flash | $0.14 | $0.42 | Baseline |
| Qwen 3.5 Flash | $0.20 | $0.80 | 43% more |
| GLM-5 | $0.30 | $1.00 | 114% more |
| Minimax M3 | $0.80 | $2.40 | 471% more |
Western Models (Premium)
| Model | Input | Output | vs DeepSeek V4 Flash |
|---|---|---|---|
| GPT-5.4 Mini | $0.15 | $0.60 | 7% more |
| Gemini 3 Flash | $0.25 | $1.00 | 79% more |
| Claude Sonnet 4 | $3.00 | $15.00 | 2,043% more |
| GPT-5.4 | $10.00 | $30.00 | 7,043% more |
| Claude Opus 4 | $15.00 | $60.00 | 10,614% more |
| GPT-5.5 | $15.00 | $60.00 | 10,614% more |
Key Takeaways for Developers
1. Chinese models dominate the value tier
Mimo and DeepSeek together offer pricing that is 10–100x cheaper than premium Western models, while maintaining competitive quality on most benchmarks.
2. Use the right model for each task
Don't use GPT-5.5 or Claude Opus 4 for simple tasks. Reserve premium models for complex reasoning and use budget models (DeepSeek V4 Flash, Mimo) for routine chat and code generation.
3. TokenPAPA lets you mix and match
With a single API key, you can route each request to the optimal model — maximum quality for critical tasks, minimum cost for everything else.
4. Context windows affect real costs
Models with larger context windows (Gemini 3 at 1M, GPT-5.5 at 2M) become more expensive with long inputs. For short queries, smaller-context models like DeepSeek V4 Flash are more cost-effective.
How to Get These Prices
All prices listed here are available through TokenPAPA — a unified API gateway that gives you access to 30+ models from a single OpenAI-compatible endpoint.
Get started in 30 seconds:
- Sign up at tokenpapa.ai — no Chinese phone required
- Get $2 free credits on signup
- Use any OpenAI SDK — just change the
base_urlandapi_key
from openai import OpenAI
client = OpenAI(
base_url="https://tokenpapa.ai/v1",
api_key="your-tokenpapa-key"
)
# Use any model — same code, different model name
response = client.chat.completions.create(
model="deepseek-v4-flash", # Change this to any model
messages=[{"role": "user", "content": "Hello!"}]
)Prices reflect TokenPAPA relay pricing as of July 27, 2026. Official provider pricing may differ. All prices in USD per 1M tokens.
このガイドはいかがですか?
