Pricing

CoreWeave Forge Inference, Sandboxes, and ARIA pricing

Ship it. Watch it. Make it better.

CoreWeave Forge hosted models

Prices shown are per 1 million tokens.

Models marked "Available for post-training" support SFT and RL. See post-training pricing for training rates.

Model
Input Tokens
Output Tokens
Cache Hit
Available for
Post-Training
Model
IBM Granite 4.2 8B
Input tokens
$0.10
Output tokens
$0.15
Cache hit
$0.05
Available for Post-Training
Model
Qwen3.8 27B
Input tokens
$0.40
Output tokens
$3.00
Cache hit
$0.15
Available for Post-Training
Model
Z.AI GLM 5.3 Flash
Input tokens
$0.15
Output tokens
$0.50
Cache hit
$0.05
Available for Post-Training
Model
DeepSeek V4-Pro-0813
Input tokens
$1.31
Output tokens
$3.96
Cache hit
$0.044
Available for Post-Training
Model
NVIDIA Nemotron 3.5 Lightning
Input tokens
$0.07
Output tokens
$0.20
Cache hit
$0.04
Available for Post-Training
Model
DeepSeek V4-Flash-0731
Input tokens
$0.13
Output tokens
$0.28
Cache hit
$0.07
Available for Post-Training
Model
Z.AI GLM 5.2
Input tokens
$0.76
Output tokens
$2.42
Cache hit
$0.14
Available for Post-Training
Model
MiniMax M3
Input tokens
$0.23
Output tokens
$0.96
Cache hit
$0.05
Available for Post-Training
Model
Moonshot AI Kimi K2.7 Code
Input tokens
$0.71
Output tokens
$3.50
Cache hit
$0.15
Available for Post-Training
Model
NVIDIA Nemotron 3 Ultra
Input tokens
$0.50
Output tokens
$2.15
Cache hit
$0.10
Available for Post-Training
Model
JetBrains Mellum2 12B A2.5B
Input tokens
$0.05
Output tokens
$0.10
Cache hit
-
Available for Post-Training
Model
IBM Granite 4.1 8B
Input tokens
$0.05
Output tokens
$0.10
Cache hit
-
Available for Post-Training
Model
DeepSeek V4-Pro
Input tokens
$1.15
Output tokens
$2.55
Cache hit
$0.20
Available for Post-Training
Model
DeepSeek V4-Flash
Input tokens
$0.14
Output tokens
$0.28
Cache hit
$0.07
Available for Post-Training
Model
Qwen3.6 27B
Input tokens
$0.60
Output tokens
$3.60
Cache hit
$0.12
Available for Post-Training
Model
Moonshot AI Kimi K2.6
Input tokens
$0.65
Output tokens
$3.41
Cache hit
$0.15
Available for Post-Training
Model
Google Gemma 4 31B
Input tokens
$0.10
Output tokens
$0.34
Cache hit
-
Available for Post-Training
Model
Qwen3.6 35B A3B
Input tokens
$0.25
Output tokens
$1.25
Cache hit
-
Available for Post-Training
Model
Qwen3.5 35B A3B
Input tokens
$0.25
Output tokens
$1.25
Cache hit
-
Available for Post-Training
Model
Deepseek V3.1
Input tokens
$0.55
Output tokens
$1.65
Cache hit
-
Available for Post-Training
Model
OpenAI GPT OSS 20B
Input tokens
$0.03
Output tokens
$0.13
Cache hit
-
Available for Post-Training
Model
OpenAI GPT OSS 120B
Input tokens
$0.03
Output tokens
$0.17
Cache hit
-
Available for Post-Training
Model
Qwen3 30B A3B
Input tokens
$0.10
Output tokens
$0.30
Cache hit
-
Available for Post-Training
Model
OpenPipe Qwen3 14B Instruct
Input tokens
$0.05
Output tokens
$0.22
Cache hit
-
Available for Post-Training
Model
Meta Llama 3.3 70B
Input tokens
$0.71
Output tokens
$0.71
Cache hit
-
Available for Post-Training
Model
Meta Llama 3.1 70B
Input tokens
$0.80
Output tokens
$0.80
Cache hit
-
Available for Post-Training
Model
Meta Llama 3.1 8B
Input tokens
$0.22
Output tokens
$0.22
Cache hit
-
Available for Post-Training

CoreWeave Sandboxes Pricing

CPU and memory pricing, metered by millisecond.

Sandboxes are isolated environments for running untrusted code. CPU and memory are billed by the millisecond on CoreWeave-managed capacity, with no internet egress fees. GPU sandboxes, available in Private Preview, add per-GPU-hour charges.

Resource
Unit
Price
Resource
CPU
Unit
per CPU-core / hour
Price
$0.128
Resource
Memory
Unit
per GiB / hour
Price
$0.0212
Resource
GPU
Unit
Private Preview
Price
Private Preview

CoreWeave ARIA Pricing

AI Research & Iteration Agent

CoreWeave ARIA is currently free for a limited time. After the promotional period, it will use consumption-based pricing. ARIA runs on OpenAI models. A single run may use multiple models (e.g., gpt-5.4-mini and gpt-5.5). All plans include an Agent Token Rate of $0.50 per million tokens. This rate applies on top of model API pricing for all token types—input, cached input, and output. The combined cost is shown below:

Model
Input / 1M
Input / 1M
Cached Input / 1M
Output / 1M
Model
gpt-5.4
Input / 1M
$3.00
Cached Input / 1M
$0.75
Output / 1M
$15.50
Model
gpt-5.4-mini
Input / 1M
$1.25
Cached Input / 1M
$0.575
Output / 1M
$5.00
Model
gpt-5.5
Input / 1M
$5.50
Cached Input / 1M
$1.00
Output / 1M
$30.50
FAQS

Frequently asked questions

How is CoreWeave ARIA priced?

How will I know how many tokens I’ve used each month?

Is CoreWeave ARIA locked to a specific model?

Will I be charged for API calls in the Playground?

How can I track and control my monthly API spend?

How much inference can I use each month?

What happens if I exceed my spending limit?

What happens when my inference credits run out?

How are CoreWeave Sandboxes priced?

How are GPU sandboxes billed?

What counts as billable sandbox time?

Do I pay for the resources I request, or the resources I use?