> ## Documentation Index
> Fetch the complete documentation index at: https://docs.polystack.tech/llms.txt
> Use this file to discover all available pages before exploring further.

# GPU and HPU Performance Optimization

> Get full accelerator performance with CPU pinning, NUMA affinity, hugepages, vendor GPU operators, GPUDirect RDMA networking, and accelerator telemetry.

## Overview

The On-Prem AI Infrastructure Platform tunes every layer between the workload and the
accelerator. Compute flavors pin CPUs and align memory and PCI devices to the same NUMA node,
Kubernetes clusters run the vendor's own accelerator operator, multi-node training runs over
RDMA with GPUDirect, and accelerator telemetry flows into Prometheus and Grafana.

<Note>
  Performance settings apply to both [GPU virtual machines](/services/ai-platform/gpu-vms-and-containers)
  and GPU node groups in Kubernetes clusters, because node groups run on the same tuned flavors.
</Note>

***

## Optimization Layers

<Tabs>
  <Tab title="Compute" icon="server">
    [Polystack Compute](/services/compute) flavors for accelerator workloads apply:

    | Setting | Effect |
    | - | - |
    | **CPU pinning** | Dedicates physical cores to the instance and removes CPU contention |
    | **NUMA affinity** | Places instance CPUs and memory on the same NUMA node |
    | **Hugepages** | Backs instance memory with large pages, reducing TLB overhead |
    | **PCI NUMA affinity** | Attaches the accelerator from the same NUMA node as the instance's CPUs and memory |
  </Tab>

  <Tab title="Kubernetes" icon="boxes-stacked">
    Each GPU node group runs its vendor's operator:

    | Operator | Accelerators | Capabilities |
    | - | - | - |
    | **NVIDIA GPU Operator** | NVIDIA GPUs | Driver and runtime management, MIG partitioning, time-slicing, DCGM monitoring |
    | **Intel Gaudi Operator** | Intel Gaudi HPUs | Gaudi driver, device plugin, and runtime management |
    | **AMD GPU Operator** | AMD GPUs | ROCm driver, device plugin, and runtime management |
  </Tab>

  <Tab title="Multi-Node Training" icon="network-wired">
    Distributed training across nodes uses:

    | Component | Role |
    | - | - |
    | **SR-IOV / RDMA networking** | High-bandwidth, low-latency interconnect between training nodes |
    | **NVIDIA Network Operator** | Manages RDMA-capable networking in the cluster and enables GPUDirect |
    | **GPUDirect** | Moves data directly between GPU memory and the network, bypassing the CPU |
  </Tab>

  <Tab title="Monitoring" icon="chart-line">
    | Exporter | Accelerators | Destination |
    | - | - | - |
    | **DCGM exporter** | NVIDIA GPUs | Prometheus and Grafana |
    | **ROCm exporter** | AMD GPUs | Prometheus and Grafana |

    Accelerator utilization, memory, temperature, and health metrics appear alongside the
    rest of the platform in [Polystack Monitoring](/services/monitoring/index).
  </Tab>
</Tabs>

***

## Data Path for Multi-Node Training

```mermaid theme={null}
graph LR
    subgraph N1[GPU Node 1]
        G1[GPU] --- NIC1[RDMA NIC<br/>SR-IOV]
    end
    subgraph N2[GPU Node 2]
        G2[GPU] --- NIC2[RDMA NIC<br/>SR-IOV]
    end
    NIC1 <-->|GPUDirect RDMA| NIC2
    G1 --> DCGM[DCGM / ROCm Exporters]
    G2 --> DCGM
    DCGM --> PROM[Prometheus] --> GRAF[Grafana]
```

***

## GPU Sharing in Kubernetes

<AccordionGroup>
  <Accordion title="MIG (Multi-Instance GPU)" icon="table-cells">
    Supported NVIDIA GPUs are partitioned into isolated GPU instances, each with dedicated
    compute and memory. The NVIDIA GPU Operator manages the MIG configuration on each node.
  </Accordion>

  <Accordion title="Time-slicing" icon="clock">
    Multiple pods share one NVIDIA GPU in turns. Time-slicing suits development and
    lightweight inference workloads that do not need a full GPU.
  </Accordion>
</AccordionGroup>

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Multi-Vendor Accelerator Support" icon="microchip" href="/services/ai-platform/multi-vendor-accelerators" color="#bf9667">
    How vendor operators map to node groups and cluster templates.
  </Card>

  <Card title="MLOps Platform" icon="diagram-project" href="/services/ai-platform/mlops" color="#bf9667">
    Distributed training with the Kubeflow Training Operator.
  </Card>

  <Card title="Hypervisor Configuration" icon="microchip" href="/services/compute/hypervisor" color="#bf9667">
    Hypervisor settings behind accelerator-optimized flavors.
  </Card>

  <Card title="Monitoring" icon="chart-line" href="/services/monitoring/index" color="#bf9667">
    Dashboards and alerting for platform and accelerator metrics.
  </Card>
</CardGroup>
