> ## Documentation Index
> Fetch the complete documentation index at: https://docs.polystack.tech/llms.txt
> Use this file to discover all available pages before exploring further.

# MLOps Platform

> Kubeflow, KServe, and MLflow come ready to use, delivered as a one-click AI cluster template with GPU node groups, the GPU Operator, and Harbor.

## Overview

The On-Prem AI Infrastructure Platform ships a complete MLOps toolchain ready to use.
**Kubeflow** covers pipelines, notebooks, hyperparameter tuning, and distributed training,
**KServe** serves models in production, and **MLflow** tracks experiments and manages models.
The whole stack is delivered as a one-click **AI cluster** template.

<Note>
  The AI cluster template is a [Polystack K8SaaS cluster template](/services/kubernetes/user-guide/cluster-templates).
  Deploying it creates a GPU-ready Kubernetes cluster with the full MLOps stack installed.
</Note>

***

## The AI Cluster Template

One template deploys all of the following:

| Layer | Included |
| - | - |
| **Compute** | GPU node group on passthrough GPU flavors |
| **Accelerator enablement** | GPU Operator |
| **ML workflow** | Kubeflow (Pipelines, Notebooks, Katib, Training Operator) |
| **Model serving** | KServe |
| **Experiment tracking and model registry** | MLflow |
| **Container registry** | Harbor |

```mermaid theme={null}
graph TD
    TPL[AI Cluster Template] --> NG[GPU Node Group]
    TPL --> GPUOP[GPU Operator]
    TPL --> KF[Kubeflow]
    TPL --> KS[KServe]
    TPL --> MLF[MLflow]
    TPL --> HB[Harbor]
    KF --> NB[Notebooks]
    KF --> PL[Pipelines]
    KF --> KT[Katib]
    KF --> TO[Training Operator]
```

***

## Toolchain

<Tabs>
  <Tab title="Kubeflow" icon="diagram-project">
    | Component | Purpose |
    | - | - |
    | **Notebooks** | GPU-backed notebook servers for interactive development |
    | **Pipelines** | Reproducible, multi-step ML workflows |
    | **Katib** | Automated hyperparameter tuning |
    | **Training Operator** | Distributed training jobs across multiple GPU nodes |
  </Tab>

  <Tab title="KServe" icon="rocket">
    KServe serves trained models as scalable inference endpoints on the cluster's GPU nodes.
  </Tab>

  <Tab title="MLflow" icon="flask">
    MLflow records experiment parameters, metrics, and artifacts, and keeps a registry of
    model versions ready for deployment.
  </Tab>

  <Tab title="Harbor" icon="box-archive">
    Harbor stores the container images for notebooks, training jobs, and model servers, with
    built-in Trivy vulnerability scanning.
  </Tab>
</Tabs>

***

## Model Lifecycle

<Steps>
  <Step title="Develop" icon="code">
    Data scientists explore data and build models in Kubeflow Notebooks on GPU nodes.
  </Step>

  <Step title="Train and tune" icon="gears">
    Kubeflow Pipelines orchestrate training, the Training Operator runs distributed jobs, and
    Katib tunes hyperparameters.
  </Step>

  <Step title="Track and register" icon="flask">
    MLflow records every run and registers the selected model version.
  </Step>

  <Step title="Serve" icon="rocket">
    KServe deploys the registered model as a production inference endpoint.
  </Step>
</Steps>

***

## Next Steps

<CardGroup cols={2}>
  <Card title="GPU and HPU Performance Optimization" icon="gauge-high" href="/services/ai-platform/performance-optimization" color="#bf9667">
    GPU sharing, multi-node training networking, and accelerator telemetry.
  </Card>

  <Card title="Runtime Vulnerability Scanning" icon="shield-halved" href="/services/ai-platform/vulnerability-scanning" color="#bf9667">
    How Harbor and StackRox secure MLOps workloads.
  </Card>

  <Card title="Deploy a Kubernetes Cluster" icon="boxes-stacked" href="/services/kubernetes/user-guide/deploy-cluster" color="#bf9667">
    Deploy clusters from Polystack K8SaaS templates.
  </Card>

  <Card title="Multi-Vendor Accelerator Support" icon="microchip" href="/services/ai-platform/multi-vendor-accelerators" color="#bf9667">
    Accelerator choices for AI clusters.
  </Card>
</CardGroup>
