InferenceDirect passes through provider inference pricing exactly as
billed — no markup on the per-token rate. Prepaid balance, no monthly fee, no surprise
invoices. Every model below is served through one OpenAI-compatible endpoint at
api.inferencedirect.com.
Anthropic
The Claude family — frontier reasoning, coding, and long-context models.
Model
Context
Input
Output
Claude Opus 4
200K
$15.00
$75.00
Claude Fable 5
200K
$11.00
$55.00
anthropic/claude-fable-5
1000K
$10.00
$50.00
Claude Opus 4.5
200K
$5.50
$27.50
Claude Opus 4.6
200K
$5.50
$27.50
Claude Opus 4.7
200K
$5.50
$27.50
Claude Opus 4.8
200K
$5.50
$27.50
Claude Opus 4.5
200K
$5.00
$25.00
anthropic/claude-opus-4-7
1000K
$5.00
$25.00
anthropic/claude-opus-4-8
1000K
$5.00
$25.00
Claude Sonnet 4.5
200K
$3.30
$16.50
Claude Sonnet 4.6
200K
$3.30
$16.50
Claude 3 Sonnet
200K
$3.00
$15.00
Claude Sonnet 4
200K
$3.00
$15.00
Claude Sonnet 4
200K
$3.00
$15.00
Claude Sonnet 4 (OpenRouter)
200K
$3.00
$15.00
Claude Sonnet 4.5
200K
$3.00
$15.00
anthropic/claude-sonnet-4-6
1000K
$3.00
$15.00
Claude Sonnet 5
200K
$2.20
$11.00
anthropic/claude-sonnet-5
1000K
$2.00
$10.00
Claude Haiku 4.5
200K
$1.10
$5.50
Claude Haiku 4.5
200K
$1.00
$5.00
anthropic/claude-haiku-4-5
200K
$1.00
$5.00
Claude 3.5 Haiku
200K
$0.80
$4.00
Claude 3 Haiku
200K
$0.25
$1.25
Claude 3 Haiku
200K
$0.25
$1.25
Amazon
Amazon's own Nova and Titan models, served natively on Bedrock.
Model
Context
Input
Output
Nova Pro
300K
$1.05
$4.20
Nova 2.0 Lite
300K
$0.43
$3.60
Nova Lite
300K
$0.078
$0.31
Nova Micro
128K
$0.046
$0.18
DeepSeek
Open-weight reasoning and general models with standout cost-efficiency — strong at code, math, and long-form reasoning.
Model
Context
Input
Output
deepseek-ai/DeepSeek-V4-Pro
1049K
$1.30
$2.60
deepseek-ai/DeepSeek-R1-0528
164K
$0.50
$2.15
deepseek-ai/DeepSeek-V3
164K
$0.32
$0.89
DeepSeek Chat (OpenRouter)
64K
$0.28
$0.88
deepseek-ai/DeepSeek-V3.1-Terminus
164K
$0.27
$0.95
deepseek-ai/DeepSeek-V3.2
164K
$0.26
$0.38
deepseek-ai/DeepSeek-V3.1
164K
$0.25
$0.95
deepseek-ai/DeepSeek-V3-0324
164K
$0.24
$0.90
deepseek-ai/DeepSeek-V4-Flash
1049K
$0.09
$0.18
Alibaba
The Qwen family — dense and MoE variants covering chat, coding, and vision.
Model
Context
Input
Output
Qwen/Qwen3.7-Max
256K
$2.50
$7.50
Qwen/Qwen3-Max
256K
$1.20
$6.00
Qwen/Qwen3-Max-Thinking
256K
$1.20
$6.00
Qwen/Qwen3.5-397B-A17B
262K
$0.45
$3.00
Qwen/Qwen2.5-72B-Instruct
33K
$0.36
$0.40
Qwen/Qwen3.6-27B
262K
$0.32
$3.20
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
262K
$0.30
$1.00
Qwen/Qwen3.5-122B-A10B
262K
$0.29
$2.40
Qwen/Qwen3.5-27B
262K
$0.26
$2.60
Qwen/Qwen3-235B-A22B-Thinking-2507
262K
$0.23
$2.30
Qwen/Qwen3-VL-235B-A22B-Instruct
262K
$0.20
$0.88
Qwen/Qwen3-VL-30B-A3B-Instruct
262K
$0.15
$0.60
Qwen/Qwen3.6-35B-A3B
262K
$0.15
$0.95
Qwen/Qwen3.5-35B-A3B
262K
$0.14
$1.00
Qwen/Qwen3-14B
41K
$0.12
$0.24
Qwen/Qwen3-30B-A3B
41K
$0.12
$0.50
Qwen/Qwen3.5-9B
262K
$0.10
$0.15
Qwen/Qwen3-235B-A22B-Instruct-2507
262K
$0.09
$0.55
Qwen/Qwen3-Next-80B-A3B-Instruct
262K
$0.09
$1.10
Qwen/Qwen3-32B
41K
$0.08
$0.28
Qwen/Qwen3-Embedding-4B
33K
$0.02
$0.00
Qwen/Qwen3-Embedding-0.6B
33K
$0.01
$0.00
Qwen/Qwen3-Embedding-8B
33K
$0.01
$0.00
Meta
Meta's open-weight Llama line — dependable general-purpose instruct models with a deep tooling ecosystem.
● Prices are per 1 million
tokens, served via AWS Bedrock, at exact provider cost — no markup.
Embedding, image, and audio models bill per provider unit and are not
listed here.
Starter
Free to start. Metered usage against a prepaid balance, personal
budgets, and the full dashboard.
Custom governance: strict reserve-then-settle enforcement,
delegated administration, fiscal-period alignment. Directory sync
is a future conversation, not a shipped feature.
provider usage at exact provider cost
= metered cost, drawn from your account's
prepaid balance. No per-token markup, ever.
Budgets cap the metered cost per member, team, and department. The
ledger records both customer price and provider cost, so finance can
reconcile to the cent.
Our revenue comes from a flat 5.5% platform fee ($0.80 minimum) charged
once when you add funds by card — 5% for crypto, no minimum. The fee is added on
top of the amount you choose and shown as its own line item at checkout, so a $10
top-up credits your balance the full $10 and charges your card $10.80 (the $0.80
minimum applies, since 5.5% of $10 is only $0.55). $10 minimum
top-up, $25,000 maximum per transaction.
Optional add-ons Metered
Enterprise Security
Cryptographic request signing (Ed25519, per tenant / user / node)
plus per-request access policy — required claims, time-of-day windows,
JSON schema validation, and mandatory signing. Prepaid:
$5 per 1,000,000 calls, billed only while enabled.
Analytics & chargeback
Usage analytics and chargeback reporting with per-user and per-team
drill-down and CSV export. Usage-based pricing drawn from your prepaid
balance; storage scales from D1 to R2 for high volume.
Add-ons are independent — fund either, both, or neither on the same prepaid balance.
Card & crypto billing Live
Top up by card (Stripe) or crypto (USDC via Coinbase) and set balance
auto-top-up from a pre-authorised allowance — pay-as-you-go from a prepaid
balance, with no subscription. Administrators can also enter manual
credit.