> ## Documentation Index
> Fetch the complete documentation index at: https://docs.polystack.tech/llms.txt
> Use this file to discover all available pages before exploring further.

# GPU-Capable VMs and Containers

> Run GPU-accelerated virtual machines with passthrough or shared GPUs, and Kubernetes clusters with dedicated GPU node groups on the same infrastructure.

## Overview

The On-Prem AI Infrastructure Platform runs GPU workloads as both **virtual machines** and
**containers**. Virtual machines receive accelerators directly from
[Polystack Compute](/services/compute), and Kubernetes clusters from
[Polystack K8SaaS](/services/kubernetes/index) place GPU workloads on dedicated GPU node groups
built on the same accelerator flavors.

<Note>
  Both models draw from the same accelerator pool. Choose virtual machines for full-OS
  workloads and dedicated appliances, and Kubernetes clusters for containerized training,
  inference, and MLOps pipelines.
</Note>

***

## Workload Models

<Tabs>
  <Tab title="Virtual Machines" icon="server">
    Polystack Compute attaches accelerators to instances in three ways:

    | Method | Vendors | Use Case |
    | - | - | - |
    | **PCI passthrough** | NVIDIA, AMD, Intel Gaudi | A full, dedicated accelerator per instance for training and high-throughput inference |
    | **vGPU (mediated devices)** | NVIDIA | One physical GPU shared across multiple instances |
    | **SR-IOV** | AMD | One physical GPU partitioned into virtual functions shared across instances |

    Each method is exposed through a dedicated instance flavor, so users request GPU capacity
    by selecting the flavor at launch.
  </Tab>

  <Tab title="Containers" icon="boxes-stacked">
    Polystack K8SaaS provisions Kubernetes clusters with **GPU node groups** whose worker
    nodes use passthrough GPU flavors. CPU-only and GPU node groups run side by side in the
    same cluster, so GPU capacity is scaled independently of general workloads.

    Clusters are provisioned with **Cluster API**, which manages the full cluster lifecycle
    declaratively: creation, scaling, node group changes, and upgrades.
  </Tab>
</Tabs>

***

## Architecture

```mermaid theme={null}
graph TD
    HW[Accelerator Hosts<br/>NVIDIA / Intel Gaudi / AMD]
    HW -->|PCI passthrough| PT[Passthrough GPU Flavors]
    HW -->|vGPU / mediated devices| VG[Shared NVIDIA GPU Flavors]
    HW -->|SR-IOV| SR[Shared AMD GPU Flavors]
    PT --> VM1[GPU Virtual Machines]
    VG --> VM2[Shared-GPU Virtual Machines]
    SR --> VM2
    PT --> NG[GPU Node Groups]
    NG --> K8S[Kubernetes Clusters<br/>provisioned with Cluster API]
    CPU[CPU Flavors] --> CNG[CPU Node Groups]
    CNG --> K8S
```

***

## Choosing Between Dedicated and Shared GPUs

<AccordionGroup>
  <Accordion title="Dedicated accelerators (PCI passthrough)" icon="microchip">
    The instance or Kubernetes node owns the entire accelerator. This delivers the full
    performance of the device and is the model used for GPU node groups in Kubernetes
    clusters. Available for NVIDIA, AMD, and Intel Gaudi accelerators.
  </Accordion>

  <Accordion title="Shared NVIDIA GPUs (vGPU)" icon="share-nodes">
    One physical NVIDIA GPU is divided into multiple virtual GPUs through mediated devices.
    Each instance receives a slice of the GPU, which increases utilization for development,
    visualization, and light inference workloads.
  </Accordion>

  <Accordion title="Shared AMD GPUs (SR-IOV)" icon="share-nodes">
    One physical AMD GPU is exposed as multiple SR-IOV virtual functions, each attached to a
    separate instance.
  </Accordion>
</AccordionGroup>

<Tip>
  Inside Kubernetes clusters, NVIDIA GPUs can be shared further with MIG and time-slicing.
  See [GPU and HPU Performance Optimization](/services/ai-platform/performance-optimization).
</Tip>

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Multi-Vendor Accelerator Support" icon="microchip" href="/services/ai-platform/multi-vendor-accelerators" color="#bf9667">
    How each vendor gets its own flavor, cluster template, and operator.
  </Card>

  <Card title="Operator Framework" icon="puzzle-piece" href="/services/ai-platform/operator-framework" color="#bf9667">
    The operator catalog installed on every GPU cluster.
  </Card>

  <Card title="Kubernetes Node Groups" icon="boxes-stacked" href="/services/kubernetes/user-guide/node-groups" color="#bf9667">
    Create and scale node groups in Polystack K8SaaS clusters.
  </Card>

  <Card title="Compute Flavors" icon="server" href="/services/compute/flavors" color="#bf9667">
    How instance flavors define compute and accelerator resources.
  </Card>
</CardGroup>
