Overview
The On-Prem AI Infrastructure Platform tunes every layer between the workload and the accelerator. Compute flavors pin CPUs and align memory and PCI devices to the same NUMA node, Kubernetes clusters run the vendor’s own accelerator operator, multi-node training runs over RDMA with GPUDirect, and accelerator telemetry flows into Prometheus and Grafana.Performance settings apply to both GPU virtual machines
and GPU node groups in Kubernetes clusters, because node groups run on the same tuned flavors.
Optimization Layers
- Compute
- Kubernetes
- Multi-Node Training
- Monitoring
Polystack Compute flavors for accelerator workloads apply:
Data Path for Multi-Node Training
GPU Sharing in Kubernetes
MIG (Multi-Instance GPU)
MIG (Multi-Instance GPU)
Supported NVIDIA GPUs are partitioned into isolated GPU instances, each with dedicated
compute and memory. The NVIDIA GPU Operator manages the MIG configuration on each node.
Time-slicing
Time-slicing
Multiple pods share one NVIDIA GPU in turns. Time-slicing suits development and
lightweight inference workloads that do not need a full GPU.
Next Steps
Multi-Vendor Accelerator Support
How vendor operators map to node groups and cluster templates.
MLOps Platform
Distributed training with the Kubeflow Training Operator.
Hypervisor Configuration
Hypervisor settings behind accelerator-optimized flavors.
Monitoring
Dashboards and alerting for platform and accelerator metrics.
