Infer dedicated compute
Dedicated GPU clusters,delivered in weeks.
Bare-metal NVIDIA capacity, reserved for you. Need GB200 NVL72 racks for a frontier run? We have them. An H100 island for a month? Sure. We stand up the cluster, hand you root, and keep it training.
8 GPUs to multi-thousand-GPU islands · terms from one month · built with NVIDIA
The fleet
The machines that matter, in stock.
Every current NVIDIA training platform, from our own racks and a vetted provider network — status updated as capacity lands.
- GPUs
- 72× Blackwell Ultra per rack
- Memory
- 288GB HBM3e per GPU
- Fabric
- NVLink 5 · Quantum-X800
- GPUs
- 72× Blackwell per rack
- Memory
- 192GB HBM3e per GPU
- Fabric
- NVLink 5 · Quantum-X800
- GPUs
- 8× B200 per node
- Memory
- 180GB HBM3e per GPU
- Fabric
- 400G Quantum-2 InfiniBand
- GPUs
- 8× H200 per node
- Memory
- 141GB HBM3e per GPU
- Fabric
- 400G Quantum-2 InfiniBand
- GPUs
- 8× H100 per node
- Memory
- 80GB HBM3 per GPU
- Fabric
- 3.2Tb/s InfiniBand per node
Need a different configuration — other node counts, storage ratios, or an earlier generation? Tell us what you're running.
Strategic partner & investorNVIDIA partnership
First in line for every new generation.
NVIDIA is an investor and our closest engineering partner. In a market where frontier GPUs are allocated long before they ship, that partnership is why the board above reads “available” — and it carries through how every cluster is designed, built, and supported.
Early allocation
Each new platform generation is allocated to us directly as it ships, so reservations are fulfilled on silicon timelines — not the resale market.
Reference architecture
Every cluster is built and burned in to NVIDIA reference designs and validated before handover.
Direct engineering support
When an issue needs NVIDIA's attention, it reaches NVIDIA engineering through the partnership — not a vendor queue.
Case study — Reka
Customer zero was our own frontier lab.
Before this platform ran a single external workload, it trained Reka's multimodal frontier models end to end. Every failure mode a large training run can hit — link flaps, straggler nodes, checkpoint stalls, silent data corruption — was found and fixed on our own runs, not yours.
That's the difference between a GPU landlord and an operator: the engineers who answer your page have debugged this exact fabric under a frontier-scale deadline of their own.
$ srun --nodes=1250 --gpus=10000 train.py
[03:14:22] step 128,400 / 300,000 · loss 1.732 · 4.1M tok/s
[03:14:31] goodput 99.2% · 9,984/10,000 GPUs healthy
[03:14:40] async checkpoint saved in 41s · next in 30m
[03:15:02] straggler node-0847 → hot spare swapped in 92s
[03:15:09] goodput 99.2% · run healthy ▌
The platform
Built for training. Tuned for goodput.
The machine is the floor — what you're buying is everything that keeps it busy.
Bare metal
No hypervisor between your job and the silicon. You get the whole machine — every FLOP you pay for.
Non-blocking InfiniBand
Rail-optimized, full-bisection fabric sized for all-reduce at cluster scale, not just node-to-node benchmarks.
Parallel storage
High-throughput parallel filesystems co-located with compute, sized so checkpoints never gate your step time.
Slurm or Kubernetes, day one
Clusters are handed over with your scheduler configured, images loaded, and a reference job already run.
Goodput SLAs
Uptime commitments plus goodput reporting — we account for the hours your job actually trained, not just powered-on time.
24/7 training engineers
Around-the-clock coverage from engineers who run frontier training, with hot spares racked and ready to swap in.
How it works
Reservation to first job, in four steps.
Typical time from signature to a running cluster: weeks.
- 01
Scope
Tell us the workload. We size the cluster, fabric, and storage with you.
- 02
Reserve
Lock capacity with a firm delivery date. One month to multi-year.
- 03
Provision
Burned in, benchmarked, your scheduler live. You get root.
- 04
Operate
24/7 NOC, goodput reporting, and a shared channel with the on-call engineer.
Straight answers
What engineers ask us first.
Blunt answers up front — the rest is a call away.
Is this bare metal or VMs?
Bare metal. Your jobs run on the machine, not beside a hypervisor — you get root, BMC access, and the full fabric.
Will nodes fail?
Yes. At cluster scale, hardware failure is a certainty, not a risk. We keep hot spares racked in the same fabric, swap failed nodes in minutes, and credit the downtime.
How fast can we be training?
Standard HGX blocks in days. Custom clusters — your scheduler, images, and storage layout — typically within weeks of signature.
What terms can we get?
One month to multi-year, priced per GPU-hour, all-in. Power, cooling, storage, fabric, and support are in the number — no ingress or egress fees.
Do you own the hardware?
Some of it. We run our own fleet and source the rest from a network of vetted providers. Either way, you contract with us: one agreement, one support channel, and the same burn-in bar and SLAs on every cluster.
Who operates the cluster?
We do — 24/7, by engineers who train frontier models on this same fleet. You get goodput reporting, so you see the hours you actually trained.
Reserve your cluster.
Tell us what you're training and when. You'll get 30 minutes with an engineer, then a sized proposal with firm availability — within 48 hours.