Pricing

LLM API Pricing — Every Model, Every Provider

Compare input and output token costs across 42 models from 9 providers. Updated July 2026.

Key Takeaways
  • GPT-4o costs $2.50/M input — 60% cheaper than GPT-4 Turbo at launch (OpenAI, 2024)
  • DeepSeek V3 is 27x cheaper than GPT-4o at $0.27/M input tokens
  • Tokonomics Pro ($49/mo) includes unlimited proxy calls and hard caps — vs Helicone Pro at $79/mo

Complete Pricing Table

Model Provider Input (per 1M tokens) Output (per 1M tokens) Cost for 1K calls*
Gemini 1.5 Flash 8B Google Gemini $0.0375 $0.1500 $0.09 Details →
Llama 3.1 8B Groq $0.0500 $0.0800 $0.07 Details →
Mistral Small Mistral AI $0.0600 $0.1800 $0.12 Details →
Gemini 1.5 Flash Google Gemini $0.0750 $0.3000 $0.18 Details →
Gemini 2.0 Flash Google Gemini $0.1000 $0.4000 $0.24 Details →
Command R Cohere $0.1500 $0.6000 $0.36 Details →
GPT-4o Mini OpenAI $0.1500 $0.6000 $0.36 Details →
Claude 3 Haiku Anthropic $0.2500 $1.2500 $0.70 Details →
Mistral 7B Mistral AI $0.2500 $0.2500 $0.30 Details →
DeepSeek Coder DeepSeek $0.2700 $1.1000 $0.66 Details →
DeepSeek V4 Flash DeepSeek $0.2700 $1.1000 $0.66 Details →
DeepSeek V3 Chat DeepSeek $0.2800 $0.4200 $0.39 Details →
DeepSeek R1 Reasoner DeepSeek $0.2800 $0.4200 $0.39 Details →
Gemini 1.0 Pro Google Gemini $0.5000 $1.5000 $1.00 Details →
Mistral Large Mistral AI $0.5000 $1.5000 $1.00 Details →
GPT-3.5 Turbo OpenAI $0.5000 $1.5000 $1.00 Details →
Llama 3.3 70B Groq $0.5900 $0.7900 $0.79 Details →
Mixtral 8x7B Mistral AI $0.7000 $0.7000 $0.84 Details →
Claude Haiku 4.5 Anthropic $1.0000 $5.0000 $2.80 Details →
Command Cohere $1.0000 $2.0000 $1.60 Details →
Codestral Mistral AI $1.0000 $3.0000 $2.00 Details →
o3 Mini OpenAI $1.1000 $4.4000 $2.64 Details →
Gemini 1.5 Pro Google Gemini $1.2500 $5.0000 $3.00 Details →
Mistral Medium Mistral AI $1.5000 $7.5000 $4.20 Details →
Claude Sonnet 5 Anthropic $2.0000 $10.0000 $5.60 Details →
Mixtral 8x22B Mistral AI $2.0000 $6.0000 $4.00 Details →
o3 OpenAI $2.0000 $8.0000 $4.80 Details →
Grok 2 xAI $2.0000 $10.0000 $5.60 Details →
Command R+ Cohere $2.5000 $10.0000 $6.00 Details →
GPT-4o OpenAI $2.5000 $10.0000 $6.00 Details →
Claude 3.5 Sonnet (Oct 2024) Anthropic $3.0000 $15.0000 $8.40 Details →
Claude Sonnet 4 Anthropic $3.0000 $15.0000 $8.40 Details →
Claude Sonnet 4.5 Anthropic $3.0000 $15.0000 $8.40 Details →
o1 Mini OpenAI $3.0000 $12.0000 $7.20 Details →
Claude Opus 4 Anthropic $5.0000 $25.0000 $14.00 Details →
Claude Opus 4.6 Anthropic $5.0000 $25.0000 $14.00 Details →
Claude 2.1 Anthropic $8.0000 $24.0000 $16.00 Details →
Claude Fable 5 Anthropic $10.0000 $50.0000 $28.00 Details →
GPT-4 Turbo OpenAI $10.0000 $30.0000 $20.00 Details →
Claude 3 Opus Anthropic $15.0000 $75.0000 $42.00 Details →
o1 OpenAI $15.0000 $60.0000 $36.00 Details →
GPT-4 OpenAI $30.0000 $60.0000 $48.00 Details →

*Cost for 1,000 API calls assuming 800 input + 400 output tokens per call.

According to Epoch AI's 2024 analysis, LLM inference costs are declining approximately 10x every 18 months — yet enterprise AI spend continues to rise as adoption scales. Tokonomics addresses this paradox by providing real-time per-call cost tracking across 60+ models at $49/month, compared to Helicone's $79/month Pro plan, making budget-first cost metering accessible to startups and SMBs.

Frequently Asked Questions

What does the Free plan include?

The Free plan includes 100 API proxy calls per month, 1 API key, 1 budget alert, basic analytics, and 30-day data retention. No credit card required.

What does the Pro plan ($49/mo) include?

Pro includes unlimited API proxy calls, 5 API keys, unlimited budget alerts, Slack/Teams notifications, hard spending caps, AI cost optimization reports, 90-day data retention, and scheduled PDF cost reports.

Is there a free trial for Pro?

There's no separate trial — the Free plan itself is a permanent free tier. You can test all core features (proxy, analytics, alerts) with 100 calls/month before upgrading to Pro.

Can I cancel anytime?

Yes. Cancel from the billing dashboard at any time. You keep Pro features until the end of your current billing period, then your account reverts to the Free plan automatically.

How does Tokonomics pricing compare to Helicone?

Helicone Pro costs $79/month and focuses on observability (traces, logs, evals). Tokonomics Pro is $49/month and focuses on budget enforcement — hard spending caps, budget alerts, and per-feature cost attribution. 38% less expensive for teams whose priority is cost control.

Does Tokonomics add latency to my API calls?

Minimal. Our proxy adds ~31ms of overhead per request (3.6% on a typical DeepSeek call), based on production benchmarks. Responses are streamed back in real-time — Tokonomics never buffers the full response.

Track your actual LLM costs in real time

Tokonomics sits between your app and any LLM provider. Every call is metered, every dollar is tracked.

Start Free →
Tokonomics

The budget-first AI cost metering proxy for any stack. Track every LLM token, set budget alerts, and never get surprised by your AI bill again.

© 2026 Tokonomics. All rights reserved.