Post-Training

CoreWeave Forge Inference, Sandboxes, and ARIA pricing
Ship it. Watch it. Make it better.
CoreWeave Forge hosted models
Prices shown are per 1 million tokens.
Models marked "Available for post-training" support SFT and RL. See post-training pricing for training rates.
Post-Training
CoreWeave Sandboxes Pricing
CPU and memory pricing, metered by millisecond.
Sandboxes are isolated environments for running untrusted code. CPU and memory are billed by the millisecond on CoreWeave-managed capacity, with no internet egress fees. GPU sandboxes, available in Private Preview, add per-GPU-hour charges.
CoreWeave ARIA Pricing
AI Research & Iteration Agent
CoreWeave ARIA is currently free for a limited time. After the promotional period, it will use consumption-based pricing. ARIA runs on OpenAI models. A single run may use multiple models (e.g., gpt-5.4-mini and gpt-5.5). All plans include an Agent Token Rate of $0.50 per million tokens. This rate applies on top of model API pricing for all token types—input, cached input, and output. The combined cost is shown below:
Frequently asked questions
How is CoreWeave ARIA priced?
ARIA is currently free for a limited time. After the promotional period, pricing will reflect the token rates of the models used.
How will I know how many tokens I’ve used each month?
A token is a mathematical representation of natural language. Log in to your account to view your billing dashboard. This dashboard will show you how many tokens you’ve used during the current and past months.
Is CoreWeave ARIA locked to a specific model?
No. ARIA can use multiple models and automatically selects the best model for each task to optimize performance. Pricing reflects the token rates of the models used.
Will I be charged for API calls in the Playground?
Yes. The CoreWeave Forge Playground runs on the same metered API infrastructure as production, so all Playground usage is billed at your current rate. You will be billed at the per-token input and output prices mentioned above.
How can I track and control my monthly API spend?
CoreWeave Forge gives you usage tracking and real-time spend dashboards so you can monitor tokens consumed and costs incurred. Log in to your account to view your billing dashboard. This dashboard will show you how many tokens you’ve used during the current and past months. You can configure spending alerts to notify organization and billing admins by email when your organization reaches one or more customizable spending thresholds expressed in USD. Learn more here
How much inference can I use each month?
Each account tier has a default spending cap to help manage costs and prevent unexpected charges. If you would like to customize your spending limit, please reach out to support. Learn more here.
What happens if I exceed my spending limit?
If you hit a spend limit, new API calls are rejected with a 429 (quota exceeded) error until your limit resets. There may be a delay in enforcing the limit, and you are responsible for any overage incurred. Learn more about the hard limits here
What happens when my inference credits run out?
How inference credits are granted depends on your plan: some tiers receive recurring monthly credits while others receive a one-time grant. Your billing dashboard is the source of truth for credits granted and consumed. Continuing to run inference after credits are exhausted requires pay-as-you-go billing, which you can enable at any time, including while credits remain. Usage is then invoiced monthly. Enterprise customers can contact their account team to enable pay-as-you-go inference if it is not enabled by default.
How are CoreWeave Sandboxes priced?
How you're billed depends on where your sandboxes run:
- Serverless, through CoreWeave Forge: Sandboxes are billed for the CPU and memory they request, at the same rates across all node types. Usage is billed through CoreWeave Forge when you authenticate with a Forge API key.
- Serverless, through CoreWeave: Same rates as above, invoiced by CoreWeave when you authenticate with a CoreWeave access token.
- On your own CKS capacity: Running sandboxes on your contracted CoreWeave Kubernetes Service (CKS) capacity costs nothing beyond the compute you already have.
Serverless usage is metered by the millisecond of runtime, and there is no charge for internet egress. Storage for filesystem snapshots and volumes is billed separately, depending on whether it uses a CoreWeave-managed bucket or your own.
How are GPU sandboxes billed?
A sandbox that requests GPUs is billed per GPU-hour on top of its CPU and memory. Unlike CPU and memory, the GPU rate depends on which GPU model you request. Contact your account team to sign up for GPU sandboxes and for GPU rates.
What counts as billable sandbox time?
Billing starts when your sandbox is ready to run commands and stops when it terminates. Scheduling, image pull, startup, and cleanup are not billed. Capturing a filesystem snapshot is billed at the same CPU and memory rates, because the sandbox is held while the snapshot is written. In between, the sandbox holds the CPU and memory it requested, so it bills whether or not it is doing any work. A sandbox terminates only when you stop it or its lifetime expires: there is no idle timeout, and your own process exiting does not stop a sandbox it started.
Do I pay for the resources I request, or the resources I use?
You pay for what you request. A sandbox declares its CPU and memory when it starts and is billed on that amount for as long as it runs, regardless of load. Right-sizing the request is the main cost lever: a sandbox that requests 8 CPU cores and uses 1 is billed for 8.