LLM API pricing comparison 2026

60 models

Compare the live price you pay for every text model available on AnyModel. Filter by provider, sort input or output rates, and calculate the cost of your own token volume before choosing an API.

Live pricing explorer

Search by model or provider, sort exact rates, then estimate a workload before you integrate.

Token cost calculator

Estimate API cost from input and output token counts. The result uses the live AnyModel rates in this table.

Estimated cost
$0.25
for this request or batch
ModelProviderAnyModel per 1M
Vendor list input / outputCheaperContext
GPT-5.6 SolPopularcx/gpt-5.6-solOpenAI$0.20×4$5.00 / $30.0025× cheaper372K
Claude Opus 5cc/claude-opus-5Anthropic$0.30×6$5.00 / $25.0016× cheaper1M
Claude Opus 4.8Popularcc/claude-opus-4-8Anthropic$0.20×4$5.00 / $25.0025× cheaper1M
GLM-5.3-Flashglm/glm-5.3-flashZhipu$0.025×0.5$0.15 / $0.505.9× cheaper1M
Claude Sonnet 5cc/claude-sonnet-5Anthropic$0.15×3$2.00 / $10.0013× cheaper1M
Kimi K3kmc/k3Moonshot$0.15×3$3.00 / $15.0019× cheaper256K
GPT-5.5Popularcx/gpt-5.5OpenAI$0.15×3$5.00 / $30.0033× cheaper400K
GPT-5.6 Lunacx/gpt-5.6-lunaOpenAI$0.075×1.5$1.00 / $6.0013× cheaper272K
Claude Opus 4.7cc/claude-opus-4-7Anthropic$0.20×4$5.00 / $25.0025× cheaper1M
GPT-5.6 TerraPopularcx/gpt-5.6-terraOpenAI$0.15×3$2.50 / $15.0016× cheaper272K
Grok 4.6Populargcli/grok-4.6xAI$0.10×2$2.00 / $6.0020× cheaper500K
Claude Sonnet 4.6Popularag/claude-sonnet-4-6Anthropic$0.12×2.4$3.00 / $15.0025× cheaper1M
GPT-5.4Popularcx/gpt-5.4OpenAI$0.10×2$2.50 / $15.0025× cheaper400K
Claude Opus 4.6cc/claude-opus-4-6Anthropic$0.20×4$5.00 / $25.0025× cheaper1M
GLM-5.3glm/glm-5.3Zhipu$0.075×1.5$1.40 / $4.4018× cheaper200K
Kimi K2.7 Codekmc/kimi-for-codingMoonshot$0.075×1.5256K
Claude Haiku 4.5cc/claude-haiku-4-5-20251001Anthropic$0.08×1.6$1.00 / $5.0012× cheaper200K
Free Models Auto Routeram/freeAnyModel$0.00×0Varies
GPT-5.4 minicx/gpt-5.4-miniOpenAI$0.075×1.5$0.75 / $4.509.9× cheaper400K
MiniMax M3am/minimax-m3MiniMax$0.00×0512K
Gemini 3.5 Flashag/gemini-3.5-flash-highGoogle$0.03×0.6$1.50 / $9.0050× cheaper1M
Grok 4.20 Multi-Agent ResearchPopularxai/grok-4.20-multi-agent-0309xAI$0.10×2$1.25 / $2.5012× cheaper1M
Gemini 3.1 Flash-Liteag/gemini-3.1-flash-lite-previewGoogle$0.03×0.6$0.25 / $1.508.3× cheaper1M
Nemotron 3 Ultraam/nemotron-3-ultra-550b-a55bNVIDIA$0.00×0256K
Qwen3.7 Maxqwen/qwen3.7-maxQwen$0.0125×0.25$2.50 / $7.50200× cheaper1M
Gemma 4 31Bam/gemma-4-31b-itGoogle$0.00×0128K
Kimi K3 1Mam/kimi-k3Moonshot$0.05×1$3.00 / $15.0060× cheaper1M
Gemini 2.5 Flashag/gemini-2.5-flashGoogle$0.03×0.61M
Gemini 2.5 Flash Liteag/gemini-2.5-flash-liteGoogle$0.03×0.61M
Gemini 2.5 Proag/gemini-2.5-proGoogle$0.085×1.71M
Gemini 3 Flashag/gemini-3-flashGoogle$0.03×0.61M
Gemini 3.6 Flashag/gemini-3.6-flash-highGoogle$0.03×0.61M
Gemini 3.7 Flashag/gemini-3.7-flash-highGoogle$0.03×0.61M
Gemini 3.1 Proag/gemini-pro-agentGoogle$0.085×1.71M
GPT-OSS 120Bag/gpt-oss-120b-mediumOpenAI$0.025×0.5128K
DiffusionGemma 26B A4B ITam/diffusiongemma-26b-a4b-itGoogle$0.00×0128K
GPT-OSS 20Bam/gpt-oss-20bOpenAI$0.00×0128K
Laguna XS 2.1am/laguna-xs-2.1Poolside$0.00×0128K
Llama 3.2 11B Vision Instructam/llama-3.2-11b-vision-instructMeta$0.00×0128K
Mistral Nemotronam/mistral-nemotronMistral AI$0.00×0128K
Nemotron 3 Nano 30B A3Bam/nemotron-3-nano-30b-a3bNVIDIA$0.00×0256K
Nemotron 3 Nano Omni 30B A3B Reasoningam/nemotron-3-nano-omni-30b-a3b-reasoningNVIDIA$0.00×0256K
Nemotron 3 Super 120B A12Bam/nemotron-3-super-120b-a12bNVIDIA$0.00×0256K
Nemotron 3.5 Content Safetyam/nemotron-3.5-content-safetyNVIDIA$0.00×032K
Nemotron 3.5 Lightning 30B A3Bam/nemotron-3.5-lightning-30b-a3bNVIDIA$0.00×0256K
Riva Translate 4B Instruct v2am/riva-translate-4b-instruct-v2NVIDIA$0.00×08K
DeepSeek V4 Flashds/deepseek-v4-flashDeepSeek$0.0025×0.051M
DeepSeek-V4-Prods/deepseek-v4-proDeepSeek$0.0075×0.151M
GLM 4.6Vglm/glm-4.6vZhipu$0.015×0.3128K
GLM 4.7glm/glm-4.7Zhipu$0.015×0.3200K
GLM 5glm/glm-5Zhipu$0.025×0.5200K
GLM-5.1glm/glm-5.1Zhipu$0.075×1.5200K
GLM-5.2glm/glm-5.2Zhipu$0.075×1.5200K
Qwen3.6 Flashqwen/qwen3.6-flashQwen$0.0125×0.251M
Qwen 3.7 Plusqwen/qwen3.7-plusQwen$0.0125×0.251M
Qwen3.8 Maxqwen/qwen3.8-maxQwen$0.0125×0.251M
Grok 4.20xai/grok-4.20-0309-reasoningxAI$0.05×11M
Grok 4.3xai/grok-4.3xAI$0.05×11M
Grok 4.5xai/grok-4.5xAI$0.10×2500K
Grok Build 0.1xai/grok-build-0.1xAI$0.025×0.5256K

AnyModel prices are shown per 1M tokens. Most models use one flat rate; models with separate input/output billing show both rates. Vendor list prices are shown only when directly comparable. No subscription, no minimums.

Image and video pricing

Media is billed per unit of output, not per 1M tokens. A video job costs its duration times the chosen model and resolution rate, up to 15 seconds; an image costs the base rate scaled by quality and size.

WhatUnitBalance tokens≈ price
Grok Imagine Video 480pper second125,000$0.0063
Grok Imagine Video 720pper second175,000$0.0087
Grok Imagine Video 1.5 480pper second200,000$0.01
Grok Imagine Video 1.5 720pper second350,000$0.0175
Grok Imagine Video 1.5 1080pper second625,000$0.0313
Image (gpt-image-2)per image, 1024×1024, medium250,000$0.0125

Example: an 8-second 480p clip costs 1M tokens. The charge is booked once, when the provider accepts the job, and polling a running job is free. A job that ends without a video is returned in full — by itself for jobs started in Video Studio, and on the poll that first reports the failure for jobs created through the API.

How to compare LLM API pricing

Input and output rates matter separately. Retrieval, classification and long-document workloads are input-heavy; agents and content generation can spend much more on output.

The AnyModel columns show current customer rates per 1 million tokens. Vendor list prices are context only and may not include caching, batch discounts, regional taxes or special tiers.

Price is only one production variable. Test response quality, latency, rate limits and availability with the same prompts before moving a critical workload.

LLM API pricing FAQ

What does price per 1M tokens mean?

It is the charge for one million input or output tokens. Your estimated cost is input tokens divided by one million times the input rate, plus output tokens divided by one million times the output rate.

Are the prices in this table live?

The AnyModel rates are resolved from the same current pricing catalog used by the service. Model availability and rates can change, so verify the page and dashboard before a production launch.

Is there a monthly API subscription?

No. AnyModel uses a prepaid pay-as-you-go balance with no monthly platform fee or minimum spend.

One key, every model, pay per token

Point any OpenAI-compatible client at AnyModel and switch models by name — no separate billing per provider. Fund one prepaid balance and pay only for what you use.