ONLINEUPTIME:99.99% Network Availability SLAONLINEPOWER:100 kW/Rack Direct-to-Chip Liquid CoolingONLINEUS-East:128 x H100 SXM AvailableONLINEUS-West:64 x H200 AvailableLIMITEDEU-Central:GB200 NVL72 Reserve OnlyFABRIC:GPUDirect RDMA Enabled · 3.2 Tbps InfiniBand · Data Egress: $0.00/GB
ONLINEUPTIME:99.99% Network Availability SLAONLINEPOWER:100 kW/Rack Direct-to-Chip Liquid CoolingONLINEUS-East:128 x H100 SXM AvailableONLINEUS-West:64 x H200 AvailableLIMITEDEU-Central:GB200 NVL72 Reserve OnlyFABRIC:GPUDirect RDMA Enabled · 3.2 Tbps InfiniBand · Data Egress: $0.00/GB

Enterprise compute

A control plane built for multi-node training, not a Docker wrapper.

kwctl handles node discovery, dynamic MIG partitioning, and pre-flight GPU health gating so a failing card never takes down a 512-GPU run.

kwctl
$ kwctl cluster create --nodes 16 --gpu b200-nvlink --network infiniband-3.2t

Bare-Metal Kubernetes, Ray & Slurm

Pre-configured Ray clusters, Slurm workload scheduling, and bare-metal Kubernetes operators tuned for multi-node distributed training.

  • Ray autoscaling groups
  • Slurm partitions & fair-share
  • K8s GPU operator + device plugin

InfiniBand & RoCE v2 Topology

Non-blocking leaf-spine fabric at 3.2 Tbps with GPUDirect RDMA so GPUs read and write remote VRAM without traversing CPU or host RAM.

  • Quantum-2 NDR InfiniBand
  • RoCE v2 lossless Ethernet
  • SHARP in-network reductions

kwctl Control Plane & API

Proprietary lightweight orchestration for automated node discovery, dynamic MIG partitioning, and pre-flight health checks that isolate failing GPUs.

  • Automated node discovery
  • MIG dynamic partitioning
  • Pre-run GPU health gating

Checkpointing & Fault Recovery

Long-run resilience with distributed checkpoint offload and automated node replacement or migration inside a 15-minute SLA window.

  • Async checkpoint offload
  • 15-min node replacement SLA
  • Job resume from last shard

Cluster architectures

Reference topologies

ArchitectureScaleFabricPowerWorkload
GB200 NVL72 Rack72 GPUs · 1 rackNVLink 5 domain + NDR IB~120 kW liquid-cooledFrontier-scale foundation model training
B200 SXM Pod16 nodes · 128 GPUs3.2 Tbps non-blocking InfiniBand100 kW/rack direct-to-chipMulti-node pre-training with RDMA
H200 Inference Fleet8 nodes · 64 GPUsRoCE v2 lossless Ethernet60 kW/rack hybrid loopHigh-volume LLM serving at low TTFT
H100 Training Cell32 nodes · 256 GPUsQuantum-2 InfiniBand leaf-spine45 kW/rackDistributed pre-training & fine-tuning

Service levels

Financially-backed 99.99% guarantees

Network fabric availability

99.99%

Automatic service credits

Node availability

99.99%

Automatic service credits

Hardware fault recovery

< 15 min

Auto replacement or migration

Sev-1 support response

< 15 min

Dedicated TAM escalation

Tenant isolation

Single-tenant

Dedicated switches + VPC peering

Encryption

At rest & in transit

Hardware-level key isolation