InferenceDirect.com Docs Pricing Status Sign in

Models — one endpoint, every vendor

Every model listed here is reachable through the same api.inferencedirect.com endpoint, with the same OpenAI-compatible request shape. No per-vendor SDKs, no separate API keys to manage — just change the model field.

Anthropic

The Claude family — frontier reasoning, coding, and long-context models.

ModelContextInputOutput
Claude Opus 4 200K $15.00 $75.00
Claude Fable 5 200K $11.00 $55.00
anthropic/claude-fable-5 1000K $10.00 $50.00
Claude Opus 4.5 200K $5.50 $27.50
Claude Opus 4.6 200K $5.50 $27.50
Claude Opus 4.7 200K $5.50 $27.50
Claude Opus 4.8 200K $5.50 $27.50
Claude Opus 4.5 200K $5.00 $25.00
anthropic/claude-opus-4-7 1000K $5.00 $25.00
anthropic/claude-opus-4-8 1000K $5.00 $25.00
Claude Sonnet 4.5 200K $3.30 $16.50
Claude Sonnet 4.6 200K $3.30 $16.50
Claude 3 Sonnet 200K $3.00 $15.00
Claude Sonnet 4 200K $3.00 $15.00
Claude Sonnet 4 200K $3.00 $15.00
Claude Sonnet 4 (OpenRouter) 200K $3.00 $15.00
Claude Sonnet 4.5 200K $3.00 $15.00
anthropic/claude-sonnet-4-6 1000K $3.00 $15.00
Claude Sonnet 5 200K $2.20 $11.00
anthropic/claude-sonnet-5 1000K $2.00 $10.00
Claude Haiku 4.5 200K $1.10 $5.50
Claude Haiku 4.5 200K $1.00 $5.00
anthropic/claude-haiku-4-5 200K $1.00 $5.00
Claude 3.5 Haiku 200K $0.80 $4.00
Claude 3 Haiku 200K $0.25 $1.25
Claude 3 Haiku 200K $0.25 $1.25

Amazon

Amazon's own Nova and Titan models, served natively on Bedrock.

ModelContextInputOutput
Nova Pro 300K $1.05 $4.20
Nova 2.0 Lite 300K $0.43 $3.60
Nova Lite 300K $0.078 $0.31
Nova Micro 128K $0.046 $0.18

DeepSeek

Open-weight reasoning and general models with standout cost-efficiency — strong at code, math, and long-form reasoning.

ModelContextInputOutput
deepseek-ai/DeepSeek-V4-Pro 1049K $1.30 $2.60
deepseek-ai/DeepSeek-R1-0528 164K $0.50 $2.15
deepseek-ai/DeepSeek-V3 164K $0.32 $0.89
DeepSeek Chat (OpenRouter) 64K $0.28 $0.88
deepseek-ai/DeepSeek-V3.1-Terminus 164K $0.27 $0.95
deepseek-ai/DeepSeek-V3.2 164K $0.26 $0.38
deepseek-ai/DeepSeek-V3.1 164K $0.25 $0.95
deepseek-ai/DeepSeek-V3-0324 164K $0.24 $0.90
deepseek-ai/DeepSeek-V4-Flash 1049K $0.09 $0.18

Alibaba

The Qwen family — dense and MoE variants covering chat, coding, and vision.

ModelContextInputOutput
Qwen/Qwen3.7-Max 256K $2.50 $7.50
Qwen/Qwen3-Max 256K $1.20 $6.00
Qwen/Qwen3-Max-Thinking 256K $1.20 $6.00
Qwen/Qwen3.5-397B-A17B 262K $0.45 $3.00
Qwen/Qwen2.5-72B-Instruct 33K $0.36 $0.40
Qwen/Qwen3.6-27B 262K $0.32 $3.20
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo 262K $0.30 $1.00
Qwen/Qwen3.5-122B-A10B 262K $0.29 $2.40
Qwen/Qwen3.5-27B 262K $0.26 $2.60
Qwen/Qwen3-235B-A22B-Thinking-2507 262K $0.23 $2.30
Qwen/Qwen3-VL-235B-A22B-Instruct 262K $0.20 $0.88
Qwen/Qwen3-VL-30B-A3B-Instruct 262K $0.15 $0.60
Qwen/Qwen3.6-35B-A3B 262K $0.15 $0.95
Qwen/Qwen3.5-35B-A3B 262K $0.14 $1.00
Qwen/Qwen3-14B 41K $0.12 $0.24
Qwen/Qwen3-30B-A3B 41K $0.12 $0.50
Qwen/Qwen3.5-9B 262K $0.10 $0.15
Qwen/Qwen3-235B-A22B-Instruct-2507 262K $0.09 $0.55
Qwen/Qwen3-Next-80B-A3B-Instruct 262K $0.09 $1.10
Qwen/Qwen3-32B 41K $0.08 $0.28
Qwen/Qwen3-Embedding-4B 33K $0.02 $0.00
Qwen/Qwen3-Embedding-0.6B 33K $0.01 $0.00
Qwen/Qwen3-Embedding-8B 33K $0.01 $0.00

Meta

Meta's open-weight Llama line — dependable general-purpose instruct models with a deep tooling ecosystem.

ModelContextInputOutput
NousResearch/Hermes-3-Llama-3.1-405B 131K $1.00 $1.00
NousResearch/Hermes-3-Llama-3.1-70B 131K $0.70 $0.70
meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo 131K $0.40 $0.40
meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 1049K $0.20 $0.80
Llama 3.2 3B 128K $0.19 $0.19
meta-llama/Llama-Guard-4-12B 164K $0.18 $0.18
Llama 3.2 1B 128K $0.13 $0.13
Llama 3.3 70B (OpenRouter) 128K $0.12 $0.30
meta-llama/Llama-3.3-70B-Instruct-Turbo 131K $0.10 $0.32
meta-llama/Llama-4-Scout-17B-16E-Instruct 328K $0.10 $0.30
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo 131K $0.02 $0.03

OpenAI

OpenAI's gpt-oss open-weight releases — Apache-licensed reasoning models in efficient MoE sizes.

ModelContextInputOutput
GPT-4o (OpenRouter) 128K $2.50 $10.00
openai/gpt-oss-120b-Turbo 131K $0.15 $0.60
openai/gpt-oss-120b 131K $0.037 $0.17
openai/gpt-oss-20b 131K $0.03 $0.14

Mistral

Efficient European open-weight models — small, fast instruct models that punch above their size.

ModelContextInputOutput
Pixtral Large 25.02 128K $2.00 $6.00
mistralai/Mistral-Small-3.2-24B-Instruct-2506 128K $0.075 $0.20
mistralai/Mistral-Small-24B-Instruct-2501 33K $0.05 $0.08
mistralai/Mistral-Nemo-Instruct-2407 131K $0.019 $0.03

NVIDIA

Nemotron models — open reasoning lines tuned for enterprise serving.

ModelContextInputOutput
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B 262K $0.50 $2.20
nvidia/Nemotron-Content-Safety-3.5 131K $0.20 $0.20
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B 262K $0.085 $0.40
nvidia/Nemotron-3-Nano-30B-A3B 262K $0.05 $0.20
nvidia/llama-nemotron-embed-vl-1b-v2 10K $0.01 $0.00

MiniMax

MoE generalist models balancing strong agentic performance with aggressive pricing.

ModelContextInputOutput
MiniMaxAI/MiniMax-M2.7-Turbo 197K $0.38 $1.70
MiniMaxAI/MiniMax-M3 524K $0.30 $1.20
MiniMaxAI/MiniMax-M2.7 197K $0.25 $1.00

Zhipu AI

The GLM family — frontier-level coding and agentic performance with fast, inexpensive Flash tiers.

ModelContextInputOutput
zai-org/GLM-5.1 203K $1.05 $3.50
zai-org/GLM-5.2 1049K $0.93 $3.00
zai-org/GLM-5 203K $0.60 $2.08
zai-org/GLM-4.6 203K $0.50 $2.00
zai-org/GLM-4.7 203K $0.40 $1.75
zai-org/GLM-4.7-Flash 203K $0.06 $0.40

Google

Gemma open models — compact, efficient instruct models.

ModelContextInputOutput
google/gemini-3.1-pro 1000K $2.00 $12.00
google/gemini-3.5-flash 1000K $1.50 $9.00
Gemini 2.5 Pro (OpenRouter) 1000K $1.25 $10.00
google/gemini-2.5-flash 1000K $0.30 $2.50
google/gemini-3.1-flash-lite 1000K $0.25 $1.50
google/gemma-4-31B-it 262K $0.13 $0.38
google/gemma-4-31B-it-turbo 262K $0.12 $0.37
google/gemma-3-27b-it 131K $0.08 $0.16
google/gemini-1.5-flash 1000K $0.075 $0.30
google/gemma-4-26B-A4B-it 262K $0.07 $0.34
google/gemma-3-12b-it 131K $0.05 $0.15
google/gemma-3-4b-it 131K $0.05 $0.10
google/gemini-1.5-flash-8b 1000K $0.037 $0.15
google/gemma-4-E4B-it 131K $0.02 $0.10
google/embeddinggemma-300m 2K $0.002 $0.00

Amazon Bedrock

Available through the unified endpoint.

ModelContextInputOutput
GPT-4o 128K $2.50 $10.00
GPT-4o (Azure) 128K $2.50 $10.00
GPT-4.1 1000K $2.00 $8.00
GPT-4.1 (Azure) 1000K $2.00 $8.00
o3 200K $2.00 $8.00
o3 (Azure) 200K $2.00 $8.00
Gemini 2.5 Pro 1000K $1.25 $10.00
o4-mini 200K $1.10 $4.40
o4-mini (Azure) 200K $1.10 $4.40
XiaomiMiMo/MiMo-V2.5-Pro 1049K $1.00 $3.00
thinkingmachines/Inkling 131K $1.00 $4.05
Sao10K/L3.1-70B-Euryale-v2.2 131K $0.85 $0.85
moonshotai/Kimi-K2.6 262K $0.75 $3.50
moonshotai/Kimi-K2.7-Code 262K $0.74 $3.50
ByteDance/Seed-2.0-code 256K $0.50 $3.00
ByteDance/Seed-2.0-pro 256K $0.50 $3.00
moonshotai/Kimi-K2.5 262K $0.45 $2.25
GPT-4.1 mini 1000K $0.40 $1.60
Gryphe/MythoMax-L2-13b 4K $0.40 $0.40
XiaomiMiMo/MiMo-V2.5 262K $0.40 $2.00
Gemini 2.5 Flash 1000K $0.30 $2.50
ByteDance/Seed-1.8 256K $0.25 $2.00
stepfun-ai/Step-3.7-Flash 262K $0.20 $1.15
GPT-4o mini 128K $0.15 $0.60
GPT-4o mini (Azure) 128K $0.15 $0.60
tencent/Hy3 262K $0.14 $0.58
ByteDance/Seed-2.0-mini 256K $0.10 $0.40
GPT-4.1 nano 1000K $0.10 $0.40
Gemini 2.0 Flash 1000K $0.10 $0.40
Gemini 2.5 Flash-Lite 1000K $0.10 $0.40
microsoft/phi-4 16K $0.07 $0.14
Sao10K/L3-8B-Lunaris-v1-Turbo 8K $0.04 $0.05
BAAI/bge-en-icl 8K $0.01 $0.00
BAAI/bge-large-en-v1.5 1K $0.01 $0.00
BAAI/bge-m3 8K $0.01 $0.00
BAAI/bge-m3-multi 8K $0.01 $0.00
intfloat/e5-large-v2 1K $0.01 $0.00
intfloat/multilingual-e5-large 1K $0.01 $0.00
intfloat/multilingual-e5-large-instruct 1K $0.01 $0.00
thenlper/gte-large 1K $0.01 $0.00
BAAI/bge-base-en-v1.5 1K $0.005 $0.00
intfloat/e5-base-v2 1K $0.005 $0.00
sentence-transformers/all-MiniLM-L12-v2 1K $0.005 $0.00
sentence-transformers/all-MiniLM-L6-v2 1K $0.005 $0.00
sentence-transformers/all-mpnet-base-v2 1K $0.005 $0.00
sentence-transformers/clip-ViT-B-32 $0.005 $0.00
sentence-transformers/clip-ViT-B-32-multilingual-v1 1K $0.005 $0.00
sentence-transformers/multi-qa-mpnet-base-dot-v1 1K $0.005 $0.00
sentence-transformers/paraphrase-MiniLM-L6-v2 1K $0.005 $0.00
shibing624/text2vec-base-chinese 1K $0.005 $0.00
thenlper/gte-base 1K $0.005 $0.00

Prices are per 1 million tokens, served via AWS Bedrock, at exact provider cost — no markup. Embedding, image, and audio models bill per provider unit and are not listed here.

Adding a model your team already has access to

Every model on this page is served on the platform's own account. If your team already has direct access to a provider — your own OpenAI, Anthropic, Azure OpenAI, Google, OpenRouter, or AWS account — you can route through it instead with Bring Your Own Key, with an optional per-key spend cap and automatic fall-back to the platform key.