Enterprise compute
A control plane built for multi-node training, not a Docker wrapper.
kwctl handles node discovery, dynamic MIG partitioning, and pre-flight GPU health gating so a failing card never takes down a 512-GPU run.
$ kwctl cluster create --nodes 16 --gpu b200-nvlink --network infiniband-3.2tBare-Metal Kubernetes, Ray & Slurm
Pre-configured Ray clusters, Slurm workload scheduling, and bare-metal Kubernetes operators tuned for multi-node distributed training.
- › Ray autoscaling groups
- › Slurm partitions & fair-share
- › K8s GPU operator + device plugin
InfiniBand & RoCE v2 Topology
Non-blocking leaf-spine fabric at 3.2 Tbps with GPUDirect RDMA so GPUs read and write remote VRAM without traversing CPU or host RAM.
- › Quantum-2 NDR InfiniBand
- › RoCE v2 lossless Ethernet
- › SHARP in-network reductions
kwctl Control Plane & API
Proprietary lightweight orchestration for automated node discovery, dynamic MIG partitioning, and pre-flight health checks that isolate failing GPUs.
- › Automated node discovery
- › MIG dynamic partitioning
- › Pre-run GPU health gating
Checkpointing & Fault Recovery
Long-run resilience with distributed checkpoint offload and automated node replacement or migration inside a 15-minute SLA window.
- › Async checkpoint offload
- › 15-min node replacement SLA
- › Job resume from last shard
Cluster architectures
Reference topologies
| Architecture | Scale | Fabric | Power | Workload |
|---|---|---|---|---|
| GB200 NVL72 Rack | 72 GPUs · 1 rack | NVLink 5 domain + NDR IB | ~120 kW liquid-cooled | Frontier-scale foundation model training |
| B200 SXM Pod | 16 nodes · 128 GPUs | 3.2 Tbps non-blocking InfiniBand | 100 kW/rack direct-to-chip | Multi-node pre-training with RDMA |
| H200 Inference Fleet | 8 nodes · 64 GPUs | RoCE v2 lossless Ethernet | 60 kW/rack hybrid loop | High-volume LLM serving at low TTFT |
| H100 Training Cell | 32 nodes · 256 GPUs | Quantum-2 InfiniBand leaf-spine | 45 kW/rack | Distributed pre-training & fine-tuning |
Service levels
Financially-backed 99.99% guarantees
Network fabric availability
99.99%
Automatic service credits
Node availability
99.99%
Automatic service credits
Hardware fault recovery
< 15 min
Auto replacement or migration
Sev-1 support response
< 15 min
Dedicated TAM escalation
Tenant isolation
Single-tenant
Dedicated switches + VPC peering
Encryption
At rest & in transit
Hardware-level key isolation
