INFERENCE THAT PERFORMS AT SCALE

Dedicated Inference

The performance, infrastructure clarity, and explicit control you need to scale

WHY DEDICATED INFERENCE

Maximize control without owning the cluster

With Dedicated Inference from CoreWeave, you bring the model and make the architectural choices that matter: GPU class, runtime, scaling, routing. CoreWeave runs the cluster, manages availability, and keeps performance and cost legible as you scale.

GPU choice that matches the workload

Pick the GPU class that fits your latency, throughput, and cost targets. Single-node or distributed multi-node serving, billed per GPU-hour.

Bring your own weights, stored in CoreWeave Object Storage

Point deployments at fine-tuned checkpoints, custom architectures, or OSS weights in CoreWeave Object Storage.

Open runtimes, OpenAI-compatible endpoints

vLLM and SGLang runtimes with OpenAI-compatible endpoints out of the box. Swap models, runtimes, or GPU classes without rebuilding the serving stack.

Gateway-managed routing and traffic control

A tenant-isolated gateway runs authentication, load balancing, and request routing across replicas. Optimize for latency, data locality, or compliance.

Cost that maps to infrastructure

Per GPU-hour billing against your chosen GPU class and capacity model. No egress fees, no ingress fees, no service markup.

WHO IT'S BUILT FOR

For teams that need execution visibility without the ops overhead

The teams that get the most value from Dedicated Inference sit between “just use an API” and “run our own Kubernetes cluster.”

AI TEAMS

Building custom and fine-tuned models

Teams deploying domain-specific or fine-tuned models who need custom weights support, GPU selection, and a path from training to production on the same infrastructure, without taking on cluster operations.

ENTERPRISE

Running compliance and isolation-sensitive workloads

Organizations that need single-tenant GPU nodes, Availability Zone placement for data sovereignty, and predictable SLA targets for customer-facing AI systems without assembling bespoke infrastructure.

PLATFORM TEAMS

Scaling from single-node to multi-node

ML platform teams who need to grow into distributed inference while preserving GPU choice, runtime flexibility, and per-GPU-hour cost visibility as workloads expand.

COREWEAVE INFERENCE PATHS

Inference on your terms

Three inference paths built on the award-winning CoreWeave Cloud. Move between them as your workloads evolve—without replatforming—so you always get predictable performance and infrastructure-aligned economics.

Serverless inference

Pay-per-token inference on a curated OSS model catalog. No clusters to manage. Built-in tracing, evals, and observability for AI applications and agents.

Best for
Rapid iteration and AI app development
Infrastructure management
Low — API only
Model support
Curated OSS + LoRAs
Pricing
Pay-per-token
Dedicated inference

Deploy custom weights on explicitly chosen GPUs. CoreWeave operates routing, scaling, and lifecycle. Full execution transparency, no cluster management.

Best for
Custom model serving at production scale
Infrastructure management
Medium — GPU, Zone, Runtime
Model support
Open-source or custom weights
Pricing
Pay-per-GPU-hour
Inference on CKS

Full self-managed inference on CoreWeave Kubernetes Service. Own the entire serving stack—runtimes, scheduling, autoscaling, multi-node topology—on dedicated bare-metal GPU nodes.

Best for
Full infrastructure ownership and deep tuningt
Infrastructure management
High — full Kubernetes control
Model support
Any
Pricing
Pay-per-GPU-hour (Reserved, On-demand, Spot, Flex)
INFERENCE IS PART OF COREWEAVE FORGE

Serve models where your training runs

CoreWeave Forge connects the AI loop and keeps you free to build with any cloud, model, or framework, so improvement compounds with every version. Traces from training, scores from evaluation, and your live inference all run on Forge, with no tool handoffs and no rebuilding experiment context by hand.

Run

Improve

Evaluate

HOW IT HELPS

What can you do with Dedicated Inference?

Serve custom models in production without cluster ownership

Bring your own weights, select your GPU, and ship to a live endpoint. CoreWeave operates everything between your model artifact and your users. No cluster setup, no Kubernetes expertise, no infra team required to maintain the serving layer.

Keep cost transparent and predictable as inference scales

Pay-per-GPU-hour billing against explicitly chosen GPU classes means cost stays tied to infrastructure decisions, not abstract consumption units. Existing CoreWeave reserved nodes can be directed to Dedicated Inference at contracted rates without incremental fees.

Move from fine-tuning to serving without replatforming

Weights stored in CoreWeave AI Object Storage deploy directly without data movement, whether you trained them on CoreWeave or brought them in. Just bring your weights to Dedicated Inference for production on the same infrastructure and runtimes.

HOW IT WORKS

Four steps to a live endpoint

Provision a gateway, configure your deployment, send requests, observe. You make the architectural choices; CoreWeave runs the cluster.

1. Create a gateway

Pick your CoreWeave Availability Zone. CoreWeave provides a tenant-isolated gateway that handles authentication, load balancing, and external routing.

2. Create a deployment

Point to model weights in CoreWeave AI Object Storage. Select your GPU type and inference runtime. Set min/max replica counts.

3. Send an inference request

The gateway exposes an OpenAI-compatible API. Your endpoint is ready to go. CoreWeave schedules, serves, and scales.

4. Monitor and iterate

Track performance, error rates, and GPU utilization in Grafana. Update configuration or swap model weights without taking the endpoint down.

Frequently asked questions

How is Dedicated Inference different from Inference on CKS?

Can I redirect existing CoreWeave reserved capacity to Dedicated Inference?

Can I update a deployed model without downtime?

What is a gateway, and do I manage it?

Where do my model weights need to live?

Which inference runtimes are supported?

Builder Resource Center: Learn from every run

Explore demos, code, and technical resources for every stage of the AI loop. Learn how researchers, developers, and CoreWeave engineers build, observe, evaluate, and improve AI systems—and put those insights to work.