Infer dedicated compute

Dedicated GPU clusters,delivered in weeks.

Bare-metal NVIDIA capacity, reserved for you. Need GB200 NVL72 racks for a frontier run? We have them. An H100 island for a month? Sure. We stand up the cluster, hand you root, and keep it training.

8 GPUs to multi-thousand-GPU islands · terms from one month · built with NVIDIA

5current NVIDIA platforms, GB300 → H100
8 → 10k+GPUs per deployment
Weeksfrom signature to first job
24/7run by frontier training engineers

The fleet

The machines that matter, in stock.

Every current NVIDIA training platform, from our own racks and a vetted provider network — status updated as capacity lands.

Rack-scale · NVL72 systems
Reservations open
GB300NVIDIA NVL72
GPUs
72× Blackwell Ultra per rack
Memory
288GB HBM3e per GPU
Fabric
NVLink 5 · Quantum-X800
Reserve GB300
Available now
GB200NVIDIA NVL72
GPUs
72× Blackwell per rack
Memory
192GB HBM3e per GPU
Fabric
NVLink 5 · Quantum-X800
Reserve GB200
Node-scale · HGX systems
B200NVIDIA HGX
Available now
GPUs
8× B200 per node
Memory
180GB HBM3e per GPU
Fabric
400G Quantum-2 InfiniBand
Reserve B200
H200NVIDIA HGX
Available now
GPUs
8× H200 per node
Memory
141GB HBM3e per GPU
Fabric
400G Quantum-2 InfiniBand
Reserve H200
H100NVIDIA HGX
Available now
GPUs
8× H100 per node
Memory
80GB HBM3 per GPU
Fabric
3.2Tb/s InfiniBand per node
Reserve H100

Need a different configuration — other node counts, storage ratios, or an earlier generation? Tell us what you're running.

NVIDIA logoStrategic partner & investor

NVIDIA partnership

First in line for every new generation.

NVIDIA is an investor and our closest engineering partner. In a market where frontier GPUs are allocated long before they ship, that partnership is why the board above reads “available” — and it carries through how every cluster is designed, built, and supported.

Early allocation

Each new platform generation is allocated to us directly as it ships, so reservations are fulfilled on silicon timelines — not the resale market.

Reference architecture

Every cluster is built and burned in to NVIDIA reference designs and validated before handover.

Direct engineering support

When an issue needs NVIDIA's attention, it reaches NVIDIA engineering through the partnership — not a vendor queue.

Case study — Reka

Customer zero was our own frontier lab.

Before this platform ran a single external workload, it trained Reka's multimodal frontier models end to end. Every failure mode a large training run can hit — link flaps, straggler nodes, checkpoint stalls, silent data corruption — was found and fixed on our own runs, not yours.

That's the difference between a GPU landlord and an operator: the engineers who answer your page have debugged this exact fabric under a frontier-scale deadline of their own.

reka-core · 1,250 nodes · rank 0

$ srun --nodes=1250 --gpus=10000 train.py

[03:14:22] step 128,400 / 300,000 · loss 1.732 · 4.1M tok/s

[03:14:31] goodput 99.2% · 9,984/10,000 GPUs healthy

[03:14:40] async checkpoint saved in 41s · next in 30m

[03:15:02] straggler node-0847 → hot spare swapped in 92s

[03:15:09] goodput 99.2% · run healthy

10,000+GPUs in a single training fabric
99%+effective cluster goodput on frontier runs
3model generations trained end-to-end

The platform

Built for training. Tuned for goodput.

The machine is the floor — what you're buying is everything that keeps it busy.

Bare metal

No hypervisor between your job and the silicon. You get the whole machine — every FLOP you pay for.

Non-blocking InfiniBand

Rail-optimized, full-bisection fabric sized for all-reduce at cluster scale, not just node-to-node benchmarks.

Parallel storage

High-throughput parallel filesystems co-located with compute, sized so checkpoints never gate your step time.

Slurm or Kubernetes, day one

Clusters are handed over with your scheduler configured, images loaded, and a reference job already run.

Goodput SLAs

Uptime commitments plus goodput reporting — we account for the hours your job actually trained, not just powered-on time.

24/7 training engineers

Around-the-clock coverage from engineers who run frontier training, with hot spares racked and ready to swap in.

How it works

Reservation to first job, in four steps.

Typical time from signature to a running cluster: weeks.

  1. 01

    Scope

    Tell us the workload. We size the cluster, fabric, and storage with you.

  2. 02

    Reserve

    Lock capacity with a firm delivery date. One month to multi-year.

  3. 03

    Provision

    Burned in, benchmarked, your scheduler live. You get root.

  4. 04

    Operate

    24/7 NOC, goodput reporting, and a shared channel with the on-call engineer.

Straight answers

What engineers ask us first.

Blunt answers up front — the rest is a call away.

Is this bare metal or VMs?

Bare metal. Your jobs run on the machine, not beside a hypervisor — you get root, BMC access, and the full fabric.

Will nodes fail?

Yes. At cluster scale, hardware failure is a certainty, not a risk. We keep hot spares racked in the same fabric, swap failed nodes in minutes, and credit the downtime.

How fast can we be training?

Standard HGX blocks in days. Custom clusters — your scheduler, images, and storage layout — typically within weeks of signature.

What terms can we get?

One month to multi-year, priced per GPU-hour, all-in. Power, cooling, storage, fabric, and support are in the number — no ingress or egress fees.

Do you own the hardware?

Some of it. We run our own fleet and source the rest from a network of vetted providers. Either way, you contract with us: one agreement, one support channel, and the same burn-in bar and SLAs on every cluster.

Who operates the cluster?

We do — 24/7, by engineers who train frontier models on this same fleet. You get goodput reporting, so you see the hours you actually trained.

Reserve your cluster.

Tell us what you're training and when. You'll get 30 minutes with an engineer, then a sized proposal with firm availability — within 48 hours.