Skip to main content

Overview

The On-Prem AI Infrastructure Platform ships a complete MLOps toolchain ready to use. Kubeflow covers pipelines, notebooks, hyperparameter tuning, and distributed training, KServe serves models in production, and MLflow tracks experiments and manages models. The whole stack is delivered as a one-click AI cluster template.
The AI cluster template is a Polystack K8SaaS cluster template. Deploying it creates a GPU-ready Kubernetes cluster with the full MLOps stack installed.

The AI Cluster Template

One template deploys all of the following:

Toolchain


Model Lifecycle

Develop

Data scientists explore data and build models in Kubeflow Notebooks on GPU nodes.

Train and tune

Kubeflow Pipelines orchestrate training, the Training Operator runs distributed jobs, and Katib tunes hyperparameters.

Track and register

MLflow records every run and registers the selected model version.

Serve

KServe deploys the registered model as a production inference endpoint.

Next Steps

GPU and HPU Performance Optimization

GPU sharing, multi-node training networking, and accelerator telemetry.

Runtime Vulnerability Scanning

How Harbor and StackRox secure MLOps workloads.

Deploy a Kubernetes Cluster

Deploy clusters from Polystack K8SaaS templates.

Multi-Vendor Accelerator Support

Accelerator choices for AI clusters.