PRICING
Pricing that matches how you run AI
Token Factory is our managed AI service: you access models through an API and pay for the tokens you use, while we run everything behind them. GPU Cloud is dedicated GPU capacity for teams that want to run the infrastructure themselves.
Token Factory
Managed AI services, priced per token
Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.
Leading models through one API endpoint
Pay per token, with no GPU management
Optimized for low-latency inference
Run inference, fine-tuning, or clusters
99.9% uptime SLA on dedicated endpoints
GPU Cloud
Dedicated capacity, from 2,000 GPUs
Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.
GB200 and GB300 today, Vera Rubin next
Multi-tenant hardware and private clusters
Bare metal, Kubernetes, or VM orchestration
Capacity planning with our engineers
Networking built for distributed training
Token Factory
Managed AI services, priced per token
Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.
Leading models through one API endpoint
Pay per token, with no GPU management
Optimized for low-latency inference
Run inference, fine-tuning, or clusters
99.9% uptime SLA on dedicated endpoints
GPU Cloud
Dedicated capacity, from 2,000 GPUs
Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.
GB200 and GB300 today, Vera Rubin next
Multi-tenant hardware and private clusters
Bare metal, Kubernetes, or VM orchestration
Capacity planning with our engineers
Networking built for distributed training
Token Factory
Managed AI services, priced per token
Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.
Leading models through one API endpoint
Pay per token, with no GPU management
Optimized for low-latency inference
Run inference, fine-tuning, or clusters
99.9% uptime SLA on dedicated endpoints
GPU Cloud
Dedicated capacity, from 2,000 GPUs
Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.
GB200 and GB300 today, Vera Rubin next
Multi-tenant hardware and private clusters
Bare metal, Kubernetes, or VM orchestration
Capacity planning with our engineers
Networking built for distributed training
Scale AI infrastructure from chip to cluster
Access dedicated GPU Cloud capacity designed for teams building the next generation of AI.
