TokenPAPATokenPAPA
利用ガイドAPIリファレンスAIアプリケーションブログ

LLM API Pricing Cheat Sheet 2026: Real Costs Across 25+ Models (Updated)

2026 LLM API pricing for every major model: DeepSeek V4, GPT-5, Claude 4, Mimo, Gemini, Qwen, GLM, Kimi, Minimax, Hunyuan. Input/output costs per 1M tokens, context windows, and real-world cost analysis.

LLM API Pricing Cheat Sheet 2026: Real Costs Across 25+ Models (Updated)

Last updated: July 27, 2026 — Prices reflect real-time rates across all major LLM providers accessible via TokenPAPA. Bookmark this page for the latest data.

Finding the right LLM API for your project means comparing pricing across a dozen providers — and the price differences are staggering. A single API call to Claude Opus 4 costs 50x more than the equivalent call to DeepSeek V4 Pro.

We maintain this cheat sheet with real, verified pricing for every major model available on TokenPAPA. All prices are per 1 million tokens (input / output).


Quick Look: Cheapest to Most Expensive

RankModelInput PriceOutput PriceBest For
1Mimo V2.5$0.08$0.24Budget batch processing
2DeepSeek V4 Flash$0.14$0.42High-volume chat & coding
3Mimo V2.5 Pro$0.12$0.36Budget with quality
4GPT-5.4 Mini$0.15$0.60Fast, affordable chat
5Qwen 3.5 Flash$0.20$0.80High-speed inference
6Gemini 3 Flash$0.25$1.00Multimodal on budget
7GLM-5$0.30$1.00Chinese-optimized tasks
8Kimi K2.6$0.50$2.00Long-context (128K+)
9Minimax M3$0.80$2.40Creative generation
10DeepSeek V4 Pro$0.28$0.84Best flagship value
11Hunyuan HY3$1.00$4.00Enterprise Chinese NLP
12Gemini 3.5 Flash$1.25$5.00Fast multimodal
13Claude Sonnet 4$3.00$15.00Balanced quality/speed
14GPT-5.4$10.00$30.00High-quality reasoning
15Claude Opus 4$15.00$60.00Best-in-class output
16GPT-5.5$15.00$60.00Frontier intelligence

Complete Pricing Table

Budget Tier (Under $1/M input)

ModelProviderInputOutputContextCategory
Mimo V2.5Xiaomi$0.08$0.24128KGeneral chat
Mimo V2.5 ProXiaomi$0.12$0.36128KGeneral chat
DeepSeek V4 FlashDeepSeek$0.14$0.42128KChat, coding
GPT-5.4 MiniOpenAI$0.15$0.60128KFast chat
Qwen 3.5 FlashAlibaba$0.20$0.80128KInference
Gemini 3 FlashGoogle$0.25$1.001MMultimodal
DeepSeek V4 ProDeepSeek$0.28$0.84128KFlagship
GLM-5Zhipu AI$0.30$1.00128KChinese tasks
Qwen 3.5 PlusAlibaba$0.40$1.60128KBalanced
GLM-5.1Zhipu AI$0.50$1.80256KAdvanced reasoning
Kimi K2.6Moonshot$0.50$2.00128KLong context
GLM-5.2Zhipu AI$0.50$2.00256KDeep reasoning
Minimax M2.5MiniMax$0.60$2.00128KCreative
Minimax M2.7MiniMax$0.70$2.20256KCreative pro
Minimax M3MiniMax$0.80$2.40128KGeneration

Mid Tier ($1–$5/M input)

ModelProviderInputOutputContextCategory
Hunyuan HY3 PreviewTencent$1.00$4.00256KEnterprise
Gemini 3.5 FlashGoogle$1.25$5.001MFast multimodal
Gemini 3 Pro ImageGoogle$2.00$8.002MMultimodal pro
Claude Sonnet 4Anthropic$3.00$15.00200KBalanced

Premium Tier ($5+/M input)

ModelProviderInputOutputContextCategory
GPT-5.4OpenAI$10.00$30.001MHigh-quality
Claude Opus 4Anthropic$15.00$60.00200KBest quality
GPT-5.5OpenAI$15.00$60.002MFrontier
Gemini 3.1 ProGoogle$10.00$40.002MEnterprise multimodal

Real-World Cost Examples

Let's put these numbers in perspective. Here's what common tasks actually cost:

Task: Generate a 500-word blog post (~700 output tokens)

ModelCost
DeepSeek V4 Flash$0.00029
GPT-5.4 Mini$0.00042
Claude Sonnet 4$0.0105
GPT-5.5$0.042

Task: Analyze a 10-page document (~15K input tokens)

ModelCost
DeepSeek V4 Pro$0.0042
Mimo V2.5 Pro$0.0018
Claude Opus 4$0.225
GPT-5.4$0.15

Task: Daily batch processing (1M input + 300K output tokens)

ModelDaily CostMonthly Cost
Mimo V2.5$0.15$4.65
DeepSeek V4 Flash$0.27$8.16
Gemini 3.5 Flash$2.75$83.25
Claude Sonnet 4$7.50$225.00

Price Comparison by Model Tier

Chinese Models (Best Value)

ModelInputOutputvs DeepSeek V4 Flash
Mimo V2.5$0.08$0.2443% cheaper
DeepSeek V4 Flash$0.14$0.42Baseline
Qwen 3.5 Flash$0.20$0.8043% more
GLM-5$0.30$1.00114% more
Minimax M3$0.80$2.40471% more

Western Models (Premium)

ModelInputOutputvs DeepSeek V4 Flash
GPT-5.4 Mini$0.15$0.607% more
Gemini 3 Flash$0.25$1.0079% more
Claude Sonnet 4$3.00$15.002,043% more
GPT-5.4$10.00$30.007,043% more
Claude Opus 4$15.00$60.0010,614% more
GPT-5.5$15.00$60.0010,614% more

Key Takeaways for Developers

1. Chinese models dominate the value tier

Mimo and DeepSeek together offer pricing that is 10–100x cheaper than premium Western models, while maintaining competitive quality on most benchmarks.

2. Use the right model for each task

Don't use GPT-5.5 or Claude Opus 4 for simple tasks. Reserve premium models for complex reasoning and use budget models (DeepSeek V4 Flash, Mimo) for routine chat and code generation.

3. TokenPAPA lets you mix and match

With a single API key, you can route each request to the optimal model — maximum quality for critical tasks, minimum cost for everything else.

4. Context windows affect real costs

Models with larger context windows (Gemini 3 at 1M, GPT-5.5 at 2M) become more expensive with long inputs. For short queries, smaller-context models like DeepSeek V4 Flash are more cost-effective.


How to Get These Prices

All prices listed here are available through TokenPAPA — a unified API gateway that gives you access to 30+ models from a single OpenAI-compatible endpoint.

Get started in 30 seconds:

  1. Sign up at tokenpapa.ai — no Chinese phone required
  2. Get $2 free credits on signup
  3. Use any OpenAI SDK — just change the base_url and api_key
from openai import OpenAI

client = OpenAI(
    base_url="https://tokenpapa.ai/v1",
    api_key="your-tokenpapa-key"
)

# Use any model — same code, different model name
response = client.chat.completions.create(
    model="deepseek-v4-flash",  # Change this to any model
    messages=[{"role": "user", "content": "Hello!"}]
)

Prices reflect TokenPAPA relay pricing as of July 27, 2026. Official provider pricing may differ. All prices in USD per 1M tokens.

このガイドはいかがですか?