Skip to main content

Overview

The On-Prem AI Infrastructure Platform runs GPU workloads as both virtual machines and containers. Virtual machines receive accelerators directly from Polystack Compute, and Kubernetes clusters from Polystack K8SaaS place GPU workloads on dedicated GPU node groups built on the same accelerator flavors.
Both models draw from the same accelerator pool. Choose virtual machines for full-OS workloads and dedicated appliances, and Kubernetes clusters for containerized training, inference, and MLOps pipelines.

Workload Models

Polystack Compute attaches accelerators to instances in three ways:Each method is exposed through a dedicated instance flavor, so users request GPU capacity by selecting the flavor at launch.

Architecture


Choosing Between Dedicated and Shared GPUs

The instance or Kubernetes node owns the entire accelerator. This delivers the full performance of the device and is the model used for GPU node groups in Kubernetes clusters. Available for NVIDIA, AMD, and Intel Gaudi accelerators.
One physical NVIDIA GPU is divided into multiple virtual GPUs through mediated devices. Each instance receives a slice of the GPU, which increases utilization for development, visualization, and light inference workloads.
One physical AMD GPU is exposed as multiple SR-IOV virtual functions, each attached to a separate instance.
Inside Kubernetes clusters, NVIDIA GPUs can be shared further with MIG and time-slicing. See GPU and HPU Performance Optimization.

Next Steps

Multi-Vendor Accelerator Support

How each vendor gets its own flavor, cluster template, and operator.

Operator Framework

The operator catalog installed on every GPU cluster.

Kubernetes Node Groups

Create and scale node groups in Polystack K8SaaS clusters.

Compute Flavors

How instance flavors define compute and accelerator resources.