InferenceDirect.com Docs Pricing Status Sign in

AI API pricing — pay per token, no subscription

InferenceDirect passes through provider inference pricing exactly as billed — no markup on the per-token rate. Prepaid balance, no monthly fee, no surprise invoices. Every model below is served through one OpenAI-compatible endpoint at api.inferencedirect.com.

Anthropic

The Claude family — frontier reasoning, coding, and long-context models.

ModelContextInputOutput
Claude Opus 4 200K $15.00 $75.00
Claude Fable 5 200K $11.00 $55.00
anthropic/claude-fable-5 1000K $10.00 $50.00
Claude Opus 4.5 200K $5.50 $27.50
Claude Opus 4.6 200K $5.50 $27.50
Claude Opus 4.7 200K $5.50 $27.50
Claude Opus 4.8 200K $5.50 $27.50
Claude Opus 4.5 200K $5.00 $25.00
anthropic/claude-opus-4-7 1000K $5.00 $25.00
anthropic/claude-opus-4-8 1000K $5.00 $25.00
Claude Sonnet 4.5 200K $3.30 $16.50
Claude Sonnet 4.6 200K $3.30 $16.50
Claude 3 Sonnet 200K $3.00 $15.00
Claude Sonnet 4 200K $3.00 $15.00
Claude Sonnet 4 200K $3.00 $15.00
Claude Sonnet 4 (OpenRouter) 200K $3.00 $15.00
Claude Sonnet 4.5 200K $3.00 $15.00
anthropic/claude-sonnet-4-6 1000K $3.00 $15.00
Claude Sonnet 5 200K $2.20 $11.00
anthropic/claude-sonnet-5 1000K $2.00 $10.00
Claude Haiku 4.5 200K $1.10 $5.50
Claude Haiku 4.5 200K $1.00 $5.00
anthropic/claude-haiku-4-5 200K $1.00 $5.00
Claude 3.5 Haiku 200K $0.80 $4.00
Claude 3 Haiku 200K $0.25 $1.25
Claude 3 Haiku 200K $0.25 $1.25

Amazon

Amazon's own Nova and Titan models, served natively on Bedrock.

ModelContextInputOutput
Nova Pro 300K $1.05 $4.20
Nova 2.0 Lite 300K $0.43 $3.60
Nova Lite 300K $0.078 $0.31
Nova Micro 128K $0.046 $0.18

DeepSeek

Open-weight reasoning and general models with standout cost-efficiency — strong at code, math, and long-form reasoning.

ModelContextInputOutput
deepseek-ai/DeepSeek-V4-Pro 1049K $1.30 $2.60
deepseek-ai/DeepSeek-R1-0528 164K $0.50 $2.15
deepseek-ai/DeepSeek-V3 164K $0.32 $0.89
DeepSeek Chat (OpenRouter) 64K $0.28 $0.88
deepseek-ai/DeepSeek-V3.1-Terminus 164K $0.27 $0.95
deepseek-ai/DeepSeek-V3.2 164K $0.26 $0.38
deepseek-ai/DeepSeek-V3.1 164K $0.25 $0.95
deepseek-ai/DeepSeek-V3-0324 164K $0.24 $0.90
deepseek-ai/DeepSeek-V4-Flash 1049K $0.09 $0.18

Alibaba

The Qwen family — dense and MoE variants covering chat, coding, and vision.

ModelContextInputOutput
Qwen/Qwen3.7-Max 256K $2.50 $7.50
Qwen/Qwen3-Max 256K $1.20 $6.00
Qwen/Qwen3-Max-Thinking 256K $1.20 $6.00
Qwen/Qwen3.5-397B-A17B 262K $0.45 $3.00
Qwen/Qwen2.5-72B-Instruct 33K $0.36 $0.40
Qwen/Qwen3.6-27B 262K $0.32 $3.20
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo 262K $0.30 $1.00
Qwen/Qwen3.5-122B-A10B 262K $0.29 $2.40
Qwen/Qwen3.5-27B 262K $0.26 $2.60
Qwen/Qwen3-235B-A22B-Thinking-2507 262K $0.23 $2.30
Qwen/Qwen3-VL-235B-A22B-Instruct 262K $0.20 $0.88
Qwen/Qwen3-VL-30B-A3B-Instruct 262K $0.15 $0.60
Qwen/Qwen3.6-35B-A3B 262K $0.15 $0.95
Qwen/Qwen3.5-35B-A3B 262K $0.14 $1.00
Qwen/Qwen3-14B 41K $0.12 $0.24
Qwen/Qwen3-30B-A3B 41K $0.12 $0.50
Qwen/Qwen3.5-9B 262K $0.10 $0.15
Qwen/Qwen3-235B-A22B-Instruct-2507 262K $0.09 $0.55
Qwen/Qwen3-Next-80B-A3B-Instruct 262K $0.09 $1.10
Qwen/Qwen3-32B 41K $0.08 $0.28
Qwen/Qwen3-Embedding-4B 33K $0.02 $0.00
Qwen/Qwen3-Embedding-0.6B 33K $0.01 $0.00
Qwen/Qwen3-Embedding-8B 33K $0.01 $0.00

Meta

Meta's open-weight Llama line — dependable general-purpose instruct models with a deep tooling ecosystem.

ModelContextInputOutput
NousResearch/Hermes-3-Llama-3.1-405B 131K $1.00 $1.00
NousResearch/Hermes-3-Llama-3.1-70B 131K $0.70 $0.70
meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo 131K $0.40 $0.40
meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 1049K $0.20 $0.80
Llama 3.2 3B 128K $0.19 $0.19
meta-llama/Llama-Guard-4-12B 164K $0.18 $0.18
Llama 3.2 1B 128K $0.13 $0.13
Llama 3.3 70B (OpenRouter) 128K $0.12 $0.30
meta-llama/Llama-3.3-70B-Instruct-Turbo 131K $0.10 $0.32
meta-llama/Llama-4-Scout-17B-16E-Instruct 328K $0.10 $0.30
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo 131K $0.02 $0.03

OpenAI

OpenAI's gpt-oss open-weight releases — Apache-licensed reasoning models in efficient MoE sizes.

ModelContextInputOutput
GPT-4o (OpenRouter) 128K $2.50 $10.00
openai/gpt-oss-120b-Turbo 131K $0.15 $0.60
openai/gpt-oss-120b 131K $0.037 $0.17
openai/gpt-oss-20b 131K $0.03 $0.14

Mistral

Efficient European open-weight models — small, fast instruct models that punch above their size.

ModelContextInputOutput
Pixtral Large 25.02 128K $2.00 $6.00
mistralai/Mistral-Small-3.2-24B-Instruct-2506 128K $0.075 $0.20
mistralai/Mistral-Small-24B-Instruct-2501 33K $0.05 $0.08
mistralai/Mistral-Nemo-Instruct-2407 131K $0.019 $0.03

NVIDIA

Nemotron models — open reasoning lines tuned for enterprise serving.

ModelContextInputOutput
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B 262K $0.50 $2.20
nvidia/Nemotron-Content-Safety-3.5 131K $0.20 $0.20
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B 262K $0.085 $0.40
nvidia/Nemotron-3-Nano-30B-A3B 262K $0.05 $0.20
nvidia/llama-nemotron-embed-vl-1b-v2 10K $0.01 $0.00

MiniMax

MoE generalist models balancing strong agentic performance with aggressive pricing.

ModelContextInputOutput
MiniMaxAI/MiniMax-M2.7-Turbo 197K $0.38 $1.70
MiniMaxAI/MiniMax-M3 524K $0.30 $1.20
MiniMaxAI/MiniMax-M2.7 197K $0.25 $1.00

Zhipu AI

The GLM family — frontier-level coding and agentic performance with fast, inexpensive Flash tiers.

ModelContextInputOutput
zai-org/GLM-5.1 203K $1.05 $3.50
zai-org/GLM-5.2 1049K $0.93 $3.00
zai-org/GLM-5 203K $0.60 $2.08
zai-org/GLM-4.6 203K $0.50 $2.00
zai-org/GLM-4.7 203K $0.40 $1.75
zai-org/GLM-4.7-Flash 203K $0.06 $0.40

Google

Gemma open models — compact, efficient instruct models.

ModelContextInputOutput
google/gemini-3.1-pro 1000K $2.00 $12.00
google/gemini-3.5-flash 1000K $1.50 $9.00
Gemini 2.5 Pro (OpenRouter) 1000K $1.25 $10.00
google/gemini-2.5-flash 1000K $0.30 $2.50
google/gemini-3.1-flash-lite 1000K $0.25 $1.50
google/gemma-4-31B-it 262K $0.13 $0.38
google/gemma-4-31B-it-turbo 262K $0.12 $0.37
google/gemma-3-27b-it 131K $0.08 $0.16
google/gemini-1.5-flash 1000K $0.075 $0.30
google/gemma-4-26B-A4B-it 262K $0.07 $0.34
google/gemma-3-12b-it 131K $0.05 $0.15
google/gemma-3-4b-it 131K $0.05 $0.10
google/gemini-1.5-flash-8b 1000K $0.037 $0.15
google/gemma-4-E4B-it 131K $0.02 $0.10
google/embeddinggemma-300m 2K $0.002 $0.00

Amazon Bedrock

Available through the unified endpoint.

ModelContextInputOutput
GPT-4o 128K $2.50 $10.00
GPT-4o (Azure) 128K $2.50 $10.00
GPT-4.1 1000K $2.00 $8.00
GPT-4.1 (Azure) 1000K $2.00 $8.00
o3 200K $2.00 $8.00
o3 (Azure) 200K $2.00 $8.00
Gemini 2.5 Pro 1000K $1.25 $10.00
o4-mini 200K $1.10 $4.40
o4-mini (Azure) 200K $1.10 $4.40
XiaomiMiMo/MiMo-V2.5-Pro 1049K $1.00 $3.00
thinkingmachines/Inkling 131K $1.00 $4.05
Sao10K/L3.1-70B-Euryale-v2.2 131K $0.85 $0.85
moonshotai/Kimi-K2.6 262K $0.75 $3.50
moonshotai/Kimi-K2.7-Code 262K $0.74 $3.50
ByteDance/Seed-2.0-code 256K $0.50 $3.00
ByteDance/Seed-2.0-pro 256K $0.50 $3.00
moonshotai/Kimi-K2.5 262K $0.45 $2.25
GPT-4.1 mini 1000K $0.40 $1.60
Gryphe/MythoMax-L2-13b 4K $0.40 $0.40
XiaomiMiMo/MiMo-V2.5 262K $0.40 $2.00
Gemini 2.5 Flash 1000K $0.30 $2.50
ByteDance/Seed-1.8 256K $0.25 $2.00
stepfun-ai/Step-3.7-Flash 262K $0.20 $1.15
GPT-4o mini 128K $0.15 $0.60
GPT-4o mini (Azure) 128K $0.15 $0.60
tencent/Hy3 262K $0.14 $0.58
ByteDance/Seed-2.0-mini 256K $0.10 $0.40
GPT-4.1 nano 1000K $0.10 $0.40
Gemini 2.0 Flash 1000K $0.10 $0.40
Gemini 2.5 Flash-Lite 1000K $0.10 $0.40
microsoft/phi-4 16K $0.07 $0.14
Sao10K/L3-8B-Lunaris-v1-Turbo 8K $0.04 $0.05
BAAI/bge-en-icl 8K $0.01 $0.00
BAAI/bge-large-en-v1.5 1K $0.01 $0.00
BAAI/bge-m3 8K $0.01 $0.00
BAAI/bge-m3-multi 8K $0.01 $0.00
intfloat/e5-large-v2 1K $0.01 $0.00
intfloat/multilingual-e5-large 1K $0.01 $0.00
intfloat/multilingual-e5-large-instruct 1K $0.01 $0.00
thenlper/gte-large 1K $0.01 $0.00
BAAI/bge-base-en-v1.5 1K $0.005 $0.00
intfloat/e5-base-v2 1K $0.005 $0.00
sentence-transformers/all-MiniLM-L12-v2 1K $0.005 $0.00
sentence-transformers/all-MiniLM-L6-v2 1K $0.005 $0.00
sentence-transformers/all-mpnet-base-v2 1K $0.005 $0.00
sentence-transformers/clip-ViT-B-32 $0.005 $0.00
sentence-transformers/clip-ViT-B-32-multilingual-v1 1K $0.005 $0.00
sentence-transformers/multi-qa-mpnet-base-dot-v1 1K $0.005 $0.00
sentence-transformers/paraphrase-MiniLM-L6-v2 1K $0.005 $0.00
shibing624/text2vec-base-chinese 1K $0.005 $0.00
thenlper/gte-base 1K $0.005 $0.00

Prices are per 1 million tokens, served via AWS Bedrock, at exact provider cost — no markup. Embedding, image, and audio models bill per provider unit and are not listed here.

Starter

Free to start. Metered usage against a prepaid balance, personal budgets, and the full dashboard.

Sign in

Business Most common

Prepaid balance plus governance: department and team budgets, warn/block enforcement, per-model policy, service keys, and chargeback exports.

Sign in

Enterprise

Custom governance: strict reserve-then-settle enforcement, delegated administration, fiscal-period alignment. Directory sync is a future conversation, not a shipped feature.

Sign in

What you pay for

provider usage at exact provider cost = metered cost, drawn from your account's prepaid balance. No per-token markup, ever.

Budgets cap the metered cost per member, team, and department. The ledger records both customer price and provider cost, so finance can reconcile to the cent.

Our revenue comes from a flat 5.5% platform fee ($0.80 minimum) charged once when you add funds by card — 5% for crypto, no minimum. The fee is added on top of the amount you choose and shown as its own line item at checkout, so a $10 top-up credits your balance the full $10 and charges your card $10.80 (the $0.80 minimum applies, since 5.5% of $10 is only $0.55). $10 minimum top-up, $25,000 maximum per transaction.

Optional add-ons Metered

Enterprise Security

Cryptographic request signing (Ed25519, per tenant / user / node) plus per-request access policy — required claims, time-of-day windows, JSON schema validation, and mandatory signing. Prepaid: $5 per 1,000,000 calls, billed only while enabled.

Analytics & chargeback

Usage analytics and chargeback reporting with per-user and per-team drill-down and CSV export. Usage-based pricing drawn from your prepaid balance; storage scales from D1 to R2 for high volume.

Add-ons are independent — fund either, both, or neither on the same prepaid balance.

Card & crypto billing Live

Top up by card (Stripe) or crypto (USDC via Coinbase) and set balance auto-top-up from a pre-authorised allowance — pay-as-you-go from a prepaid balance, with no subscription. Administrators can also enter manual credit.