PRICING

Pricing that matches how you run AI

Token Factory is our managed AI service: you access models through an API and pay for the tokens you use, while we run everything behind them. GPU Cloud is dedicated GPU capacity for teams that want to run the infrastructure themselves.

Token Factory

Managed AI services, priced per token

Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.

Leading models through one API endpoint

Pay per token, with no GPU management

Optimized for low-latency inference

Run inference, fine-tuning, or clusters

99.9% uptime SLA on dedicated endpoints

GPU Cloud

Dedicated capacity, from 2,000 GPUs

Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.

GB200 and GB300 today, Vera Rubin next

Multi-tenant hardware and private clusters

Bare metal, Kubernetes, or VM orchestration

Capacity planning with our engineers

Networking built for distributed training

Token Factory

Managed AI services, priced per token

Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.

Leading models through one API endpoint

Pay per token, with no GPU management

Optimized for low-latency inference

Run inference, fine-tuning, or clusters

99.9% uptime SLA on dedicated endpoints

GPU Cloud

Dedicated capacity, from 2,000 GPUs

Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.

GB200 and GB300 today, Vera Rubin next

Multi-tenant hardware and private clusters

Bare metal, Kubernetes, or VM orchestration

Capacity planning with our engineers

Networking built for distributed training

Token Factory

Managed AI services, priced per token

Production-ready AI models served through a simple API. We run the models, GPUs, and scaling behind it, and you pay for the tokens you use. Also available as fine-tuning and dedicated clusters.

Leading models through one API endpoint

Pay per token, with no GPU management

Optimized for low-latency inference

Run inference, fine-tuning, or clusters

99.9% uptime SLA on dedicated endpoints

GPU Cloud

Dedicated capacity, from 2,000 GPUs

Reserved, multi-tenant NVIDIA clusters for training, fine-tuning, and inference. Consume your capacity as bare metal, Kubernetes, or virtual machines, with full control of your stack.

GB200 and GB300 today, Vera Rubin next

Multi-tenant hardware and private clusters

Bare metal, Kubernetes, or VM orchestration

Capacity planning with our engineers

Networking built for distributed training

Scale AI infrastructure from chip to cluster

Access dedicated GPU Cloud capacity designed for teams building the next generation of AI.