Compare input and output token costs across 42 models from 9 providers. Updated July 2026.
| Model | Provider | Input (per 1M tokens) | Output (per 1M tokens) | Cost for 1K calls* | |
|---|---|---|---|---|---|
| Gemini 1.5 Flash 8B | Google Gemini | $0.0375 | $0.1500 | $0.09 | Details → |
| Llama 3.1 8B | Groq | $0.0500 | $0.0800 | $0.07 | Details → |
| Mistral Small | Mistral AI | $0.0600 | $0.1800 | $0.12 | Details → |
| Gemini 1.5 Flash | Google Gemini | $0.0750 | $0.3000 | $0.18 | Details → |
| Gemini 2.0 Flash | Google Gemini | $0.1000 | $0.4000 | $0.24 | Details → |
| Command R | Cohere | $0.1500 | $0.6000 | $0.36 | Details → |
| GPT-4o Mini | OpenAI | $0.1500 | $0.6000 | $0.36 | Details → |
| Claude 3 Haiku | Anthropic | $0.2500 | $1.2500 | $0.70 | Details → |
| Mistral 7B | Mistral AI | $0.2500 | $0.2500 | $0.30 | Details → |
| DeepSeek Coder | DeepSeek | $0.2700 | $1.1000 | $0.66 | Details → |
| DeepSeek V4 Flash | DeepSeek | $0.2700 | $1.1000 | $0.66 | Details → |
| DeepSeek V3 Chat | DeepSeek | $0.2800 | $0.4200 | $0.39 | Details → |
| DeepSeek R1 Reasoner | DeepSeek | $0.2800 | $0.4200 | $0.39 | Details → |
| Gemini 1.0 Pro | Google Gemini | $0.5000 | $1.5000 | $1.00 | Details → |
| Mistral Large | Mistral AI | $0.5000 | $1.5000 | $1.00 | Details → |
| GPT-3.5 Turbo | OpenAI | $0.5000 | $1.5000 | $1.00 | Details → |
| Llama 3.3 70B | Groq | $0.5900 | $0.7900 | $0.79 | Details → |
| Mixtral 8x7B | Mistral AI | $0.7000 | $0.7000 | $0.84 | Details → |
| Claude Haiku 4.5 | Anthropic | $1.0000 | $5.0000 | $2.80 | Details → |
| Command | Cohere | $1.0000 | $2.0000 | $1.60 | Details → |
| Codestral | Mistral AI | $1.0000 | $3.0000 | $2.00 | Details → |
| o3 Mini | OpenAI | $1.1000 | $4.4000 | $2.64 | Details → |
| Gemini 1.5 Pro | Google Gemini | $1.2500 | $5.0000 | $3.00 | Details → |
| Mistral Medium | Mistral AI | $1.5000 | $7.5000 | $4.20 | Details → |
| Claude Sonnet 5 | Anthropic | $2.0000 | $10.0000 | $5.60 | Details → |
| Mixtral 8x22B | Mistral AI | $2.0000 | $6.0000 | $4.00 | Details → |
| o3 | OpenAI | $2.0000 | $8.0000 | $4.80 | Details → |
| Grok 2 | xAI | $2.0000 | $10.0000 | $5.60 | Details → |
| Command R+ | Cohere | $2.5000 | $10.0000 | $6.00 | Details → |
| GPT-4o | OpenAI | $2.5000 | $10.0000 | $6.00 | Details → |
| Claude 3.5 Sonnet (Oct 2024) | Anthropic | $3.0000 | $15.0000 | $8.40 | Details → |
| Claude Sonnet 4 | Anthropic | $3.0000 | $15.0000 | $8.40 | Details → |
| Claude Sonnet 4.5 | Anthropic | $3.0000 | $15.0000 | $8.40 | Details → |
| o1 Mini | OpenAI | $3.0000 | $12.0000 | $7.20 | Details → |
| Claude Opus 4 | Anthropic | $5.0000 | $25.0000 | $14.00 | Details → |
| Claude Opus 4.6 | Anthropic | $5.0000 | $25.0000 | $14.00 | Details → |
| Claude 2.1 | Anthropic | $8.0000 | $24.0000 | $16.00 | Details → |
| Claude Fable 5 | Anthropic | $10.0000 | $50.0000 | $28.00 | Details → |
| GPT-4 Turbo | OpenAI | $10.0000 | $30.0000 | $20.00 | Details → |
| Claude 3 Opus | Anthropic | $15.0000 | $75.0000 | $42.00 | Details → |
| o1 | OpenAI | $15.0000 | $60.0000 | $36.00 | Details → |
| GPT-4 | OpenAI | $30.0000 | $60.0000 | $48.00 | Details → |
*Cost for 1,000 API calls assuming 800 input + 400 output tokens per call.
According to Epoch AI's 2024 analysis, LLM inference costs are declining approximately 10x every 18 months — yet enterprise AI spend continues to rise as adoption scales. Tokonomics addresses this paradox by providing real-time per-call cost tracking across 60+ models at $49/month, compared to Helicone's $79/month Pro plan, making budget-first cost metering accessible to startups and SMBs.
The Free plan includes 100 API proxy calls per month, 1 API key, 1 budget alert, basic analytics, and 30-day data retention. No credit card required.
Pro includes unlimited API proxy calls, 5 API keys, unlimited budget alerts, Slack/Teams notifications, hard spending caps, AI cost optimization reports, 90-day data retention, and scheduled PDF cost reports.
There's no separate trial — the Free plan itself is a permanent free tier. You can test all core features (proxy, analytics, alerts) with 100 calls/month before upgrading to Pro.
Yes. Cancel from the billing dashboard at any time. You keep Pro features until the end of your current billing period, then your account reverts to the Free plan automatically.
Helicone Pro costs $79/month and focuses on observability (traces, logs, evals). Tokonomics Pro is $49/month and focuses on budget enforcement — hard spending caps, budget alerts, and per-feature cost attribution. 38% less expensive for teams whose priority is cost control.
Minimal. Our proxy adds ~31ms of overhead per request (3.6% on a typical DeepSeek call), based on production benchmarks. Responses are streamed back in real-time — Tokonomics never buffers the full response.
Tokonomics sits between your app and any LLM provider. Every call is metered, every dollar is tracked.
Start Free →The budget-first AI cost metering proxy for any stack. Track every LLM token, set budget alerts, and never get surprised by your AI bill again.