Skip to main content

Overview

The On-Prem AI Infrastructure Platform tunes every layer between the workload and the accelerator. Compute flavors pin CPUs and align memory and PCI devices to the same NUMA node, Kubernetes clusters run the vendor’s own accelerator operator, multi-node training runs over RDMA with GPUDirect, and accelerator telemetry flows into Prometheus and Grafana.
Performance settings apply to both GPU virtual machines and GPU node groups in Kubernetes clusters, because node groups run on the same tuned flavors.

Optimization Layers

Polystack Compute flavors for accelerator workloads apply:

Data Path for Multi-Node Training


GPU Sharing in Kubernetes

Supported NVIDIA GPUs are partitioned into isolated GPU instances, each with dedicated compute and memory. The NVIDIA GPU Operator manages the MIG configuration on each node.
Multiple pods share one NVIDIA GPU in turns. Time-slicing suits development and lightweight inference workloads that do not need a full GPU.

Next Steps

Multi-Vendor Accelerator Support

How vendor operators map to node groups and cluster templates.

MLOps Platform

Distributed training with the Kubeflow Training Operator.

Hypervisor Configuration

Hypervisor settings behind accelerator-optimized flavors.

Monitoring

Dashboards and alerting for platform and accelerator metrics.