Installing the RBLN NPU Operator¶
This page covers how to perform a fresh installation of the RBLN NPU Operator using its Helm chart, verify the installation, run a quick start NPU workload, configure scheduling for specific products, and customize the chart.
For an architectural overview of the operator, see RBLN NPU Operator. For upgrade and removal, see Upgrading the NPU Operator and Uninstalling the NPU Operator.
Prerequisites¶
- Kubernetes 1.19 or later cluster with access to
kubectlandhelm - Helm 3.8 or later (required for OCI registry support)
- Node Feature Discovery installed in the cluster. To keep its lifecycle separate from the operator, we recommend deploying NFD separately using the upstream Helm chart into a dedicated
node-feature-discoverynamespace. The chart's bundlednfd.enabled=trueoption is convenient for quick trials, but it ties NFD's lifecycle to this Helm release. - A dedicated namespace for the operator, such as
rbln-system - Worker nodes equipped with NPUs
If you do not have Helm installed (or your version is below 3.8), install it first:
Creating an Image Pull Secret¶
The driver container image and rbln-smd container image are hosted on repo.rebellions.ai, which requires RBLN Portal account authentication. Before installing, create a Docker registry secret in the operator namespace:
Using this exact secret name matches the chart's default driver.imagePullSecrets, so you do not need to override any Helm values.
Driver Installation¶
Before deploying the operator chart, decide whether the operator should install the kernel driver through a container or detect a driver installed directly on each host. See NPU Driver Installation for the two modes and how to configure them with the driver.enabled chart value.
Installing the NPU Operator from the OCI Registry¶
The chart is published to Docker Hub as an OCI artifact at oci://docker.io/rebellions/rbln-npu-operator-chart. Pin a version for reproducible installation. Available versions are listed on the chart page on Docker Hub.
Override individual values with --set:
Or supply a custom values file:
Verifying the Installation¶
After Helm reports a successful installation, ensure that the RBLNClusterPolicy exists and has been reconciled. Both CRDs are scoped to the cluster, so the -n flag is not needed:
A CONTAINER value of ready means every component covering container workloads is up. A VM-PASSTHROUGH value of empty means no NPU node claims the vm-passthrough workload. See Workload Labeling by Node to opt nodes in.
If driver management is enabled, confirm that the RBLNDriver custom resource is created:
READY and DESIRED are summed across every per-node-pool DaemonSet, so they should match once every NPU node has finished installing the driver.
For scripted readiness checks, query .status.state directly:
To inspect readiness per workload type (container, vm-passthrough):
Next, check that the controller and component pods are running in the operator namespace:
If all pods show Running, the operator is healthy. Pod names follow the rbln-<component>-* pattern. See Core Components for the role of each component. The driver pod's suffix (<family>-<os>-<kernel>) identifies its node pool. See Driver Image Selection for the rules. The rbln-driver-smd pods run RSMD, which ships with each RBLNDriver.
If any pod is stuck or CrashLoopBackOff, check its logs (kubectl logs <pod> -n rbln-system) and review the Helm values for missing prerequisites.
When the operator installs the driver through the driver container (driver.enabled=true in the chart), you can verify that the kernel module is loaded and matches the requested driver version. See Verifying the running driver.
Checking NPU Status¶
Check the NPU capacity that Kubernetes reports for each node. This confirms that the device plugin exposed the expected resources. With the default devicePlugin.useGenericResourceName: true setting, the operator exposes the generic resource name rebellions.ai/npu:
If the cluster contains multiple NPU products and you need to schedule workloads onto a specific one, combine the generic rebellions.ai/npu resource with the product label applied by NPU Feature Discovery. See Targeting a Specific NPU Product below.
Creating a Pod with an NPU¶
-
Create a manifest (for example,
npu-demo-pod.yaml). This example requests fourrebellions.ai/npudevices: -
Create the Pod.
-
Verify the Pod status and resource assignment.
Usekubectl describe pod npu-podto confirm that the requested NPU resources are bound and that the Pod is scheduled onto a node with an NPU.
Targeting a Specific NPU Product¶
In a cluster with mixed NPU products (for example, both RBLN-CA25 and RBLN-CR03 cards), a Pod that requests the generic rebellions.ai/npu resource may be scheduled onto any node with an NPU. Use the product label applied by NPU Feature Discovery, rebellions.ai/npu.product, together with nodeSelector or nodeAffinity to pin the workload to a specific product.
nodeSelector: single product¶
nodeAffinity: multiple products¶
The rebellions.ai/npu.family label groups products by family; its value is the lowercase family name. Check the value on your nodes with kubectl get nodes -L rebellions.ai/npu.family before writing a selector:
Configuration Reference¶
The tables below summarize the key configuration options in values.yaml. Each section maps to an entry in the Core Components overview.
Pass any of these values with --set <key>=<value> (or -f my-values.yaml) when running helm install or helm upgrade.
To see the chart's full values list and template comments, including keys not shown in the tables below, run:
The image.* blocks follow the standard Helm chart convention with four sub-keys:
In the Default column below, this is condensed as <registry>/<repository>:<tag> plus the pullPolicy: on a second line.
Chart-wide¶
| Key | Description | Default |
|---|---|---|
nameOverride |
Prefix applied to every child resource name. Override to avoid collisions when running multiple operator instances. | rbln-npu-operator |
workloadType |
Workload mode for the cluster. Set to vm-passthrough for KubeVirt deployments. In this mode, the CRD also validates that vfioManager.enabled and sandboxDevicePlugin.enabled are true. |
container |
nfd.enabled |
Whether the chart deploys Node Feature Discovery as a subchart. Leave false and install NFD separately through its upstream Helm chart to keep its lifecycle separate from the operator. Set true only for quick trials. |
false |
podDefaults.labels |
Labels applied to every DaemonSet pod managed by the operator. | {} |
podDefaults.annotations |
Annotations applied to every DaemonSet pod managed by the operator. | {} |
podDefaults.tolerations |
Tolerations applied to every DaemonSet pod managed by the operator. | [] |
podDefaults.priorityClassName |
Priority class applied to every operator-managed DaemonSet pod. | "" |
CRD Upgrade Hook¶
| Key | Description | Default |
|---|---|---|
crds.upgrade.enabled |
Whether the chart uses a pre-install/pre-upgrade hook Job to synchronize the operator CRDs on helm install/helm upgrade. Disable it when a GitOps tool reconciles the CRDs. See Upgrading the NPU Operator. |
true |
crds.upgrade.imagePullSecrets |
Image pull secret(s) for the hook Job that runs the operator image. | [] |
crds.upgrade.tolerations |
Tolerations for the hook Job pod. | [] |
crds.upgrade.resources |
CPU/memory requests and limits for the hook Job pod. | {} |
crds.upgrade.logging.level |
Log level for the hook Job: error, warning, info, or debug. The hook runs before the operator exists, so it cannot inherit operator.logging.*. |
info |
crds.upgrade.logging.format |
Log format for the hook Job: json or text. |
json |
Operator¶
| Key | Description | Default |
|---|---|---|
operator.image.* |
Image of the controller-manager pod. Override tag to pin a specific operator version. |
docker.io/rebellions/rbln-npu-operator:<chart default>pullPolicy: IfNotPresent |
operator.replicas |
Number of operator pods. Increase to 2+ for high availability. |
1 |
operator.resources.requests |
Minimum guaranteed resources for the operator pod. | cpu: 50m, memory: 128Mi |
operator.resources.limits |
Maximum resources for the operator pod. | cpu: 500m, memory: 256Mi |
operator.metrics.enabled |
Whether to enable the authenticated operator self-metrics endpoint and its Service. | true |
operator.metrics.port |
Metrics HTTPS port. | 8443 |
operator.metrics.serviceMonitor.* |
Whether to create a ServiceMonitor for the Prometheus Operator, plus scrape options (interval, scrapeTimeout, and so on). See Operator Observability. |
enabled: false |
operator.logging.level |
Verbosity of the operator's own log stream: error, info, debug, panic, or an integer N for logr levels up to V(N). An invalid value stops the operator at startup. |
info |
operator.logging.encoder |
Log encoding: json or console. |
json |
operator.logging.timeEncoding |
Timestamp format: epoch, millis, nanos, iso8601, rfc3339, or rfc3339nano. |
rfc3339nano |
operator.logging.develMode |
Whether to enable development behavior, which panics on DPanic and dumps full Kubernetes objects into log values. Keep this set to false except when debugging. |
false |
operator.driverImageCheck |
Whether the operator verifies each composed driver image in its registry before rolling out the pool that uses it. See Registry check before rollout. | true |
operator.securityContext.runAsNonRoot |
Pod-level security context. Adjust based on your cluster's security policy. | true |
operator.affinity |
Affinity for the operator pod (e.g., to pin to control-plane nodes). | {} |
operator.tolerations |
Tolerations for the operator pod (for example, to allow scheduling on nodes with taints). | [] |
Driver Manager¶
| Key | Description | Default |
|---|---|---|
driver.enabled |
Whether the operator installs and manages the NPU driver. Leave false if drivers are already installed on hosts. |
false |
driver.image.* |
Driver container image. Override this value to pin a private mirror or a specific driver release. Leave the NPU family out of repository; the operator inserts it per node pool. |
repo.rebellions.ai/rebellions/rbln-driver:<chart default>pullPolicy: IfNotPresent |
driver.imagePullSecrets |
Image pull secret for the driver image. Change this value to match your actual secret name if you do not use the drivercred name created earlier. |
[drivercred] |
driver.nodeSelector |
Restrict driver pods to specific nodes. The selector cannot use the operator-managed keys rebellions.ai/npu.driver.owner and rebellions.ai/npu.deploy.driver. |
{} |
driver.tolerations |
Tolerations for driver pods. | [] |
driver.annotations |
Annotations for driver pods. | {} |
driver.priorityClassName |
Pod priority class. The CRD defaults to system-node-critical when this value is empty. |
"" |
driver.resources |
CPU/memory requests and limits for the driver pod. The CRD field is required; the chart fills defaults when unset. | {} |
driver.env |
Environment variables passed to the driver container (e.g., log levels). | [] |
driver.manager.image.* |
Image of the driver manager initContainer that performs reconciliation on the node. Keep the chart default: this operator requires v0.2.2 or later. |
docker.io/rebellions/rbln-k8s-driver-manager:<chart default>pullPolicy: IfNotPresent |
driver.smd.image.registrydriver.smd.image.repository |
Registry and repository of the RSMD node daemon that ships with each driver. There is no tag key: the tag always follows the driver version. | repo.rebellions.airebellions/rbln-smd |
driver.deployDefault |
Whether to deploy the default RBLNDriver that covers every node that no instance claims. Setting false removes the fallback driver from those nodes. |
true |
driver.instances |
Additional RBLNDriver resources for running per-node-group driver versions. See Multiple Driver Versions. |
{} |
driver.upgradePolicy.* keys (autoUpgrade, drain, reboot, etc.) are documented in NPU Driver Upgrade Workflow.
Device Plugin¶
| Key | Description | Default |
|---|---|---|
devicePlugin.enabled |
Whether the standard container Device Plugin is deployed. Disable when running only DRA (draKubeletPlugin) or only VM workloads. |
true |
devicePlugin.image.* |
Device Plugin image. | docker.io/rebellions/k8s-device-plugin:<chart default>pullPolicy: IfNotPresent |
devicePlugin.useGenericResourceName |
Whether to expose the generic rebellions.ai/npu resource. Keep true; setting false selects a naming mode by card that is not recommended for new deployments. |
true |
devicePlugin.otlpEndpoint |
OTLP gRPC endpoint the Device Plugin exports NPU allocation traces to (host:port or a URL including a scheme). Setting it enables tracing. |
"" |
DRA Driver¶
| Key | Description | Default |
|---|---|---|
draKubeletPlugin.enabled |
Enable on Kubernetes 1.34 or later to use Dynamic Resource Allocation. Mutually exclusive with devicePlugin.enabled. |
false |
draKubeletPlugin.image.* |
DRA kubelet plugin image. | docker.io/rebellions/k8s-dra-driver-npu:<chart default>pullPolicy: IfNotPresent |
draKubeletPlugin.driverName |
Driver name; must match the value referenced from DeviceClass.spec.config.driver. |
npu.rebellions.ai |
draKubeletPlugin.kubeletRegistrarDirectoryPath |
Host path where the plugin registers with the kubelet. | /var/lib/kubelet/plugins_registry |
draKubeletPlugin.kubeletPluginsDirectoryPath |
Host path where the plugin sockets live. | /var/lib/kubelet/plugins |
draKubeletPlugin.healthcheckPort |
TCP port for the plugin's health-check endpoint. | 51515 |
See NPU DRA Driver for full DRA usage.
Sandbox Device Plugin¶
| Key | Description | Default |
|---|---|---|
sandboxDevicePlugin.enabled |
Whether the VFIO Sandbox Device Plugin is deployed. Enable for KubeVirt or other VM environments. | false |
sandboxDevicePlugin.image.* |
Sandbox Device Plugin image. | docker.io/rebellions/k8s-device-plugin:<chart default>pullPolicy: IfNotPresent |
Resource names are discovered automatically per NPU product. There is no resource list to configure. See Resource naming.
VFIO Manager¶
| Key | Description | Default |
|---|---|---|
vfioManager.enabled |
Whether the VFIO bind/unbind helper is deployed. Enable it together with the Sandbox Device Plugin for VM passthrough. | false |
vfioManager.image.* |
VFIO Manager image. | docker.io/rebellions/rbln-vfio-manager:<chart default>pullPolicy: IfNotPresent |
vfioManager.driverManager.image.* |
Driver manager image used by the VFIO Manager initContainer. | docker.io/rebellions/rbln-k8s-driver-manager:<chart default>pullPolicy: IfNotPresent |
vfioManager.driverManager.env |
Environment variables for that initContainer. | [] |
Container Toolkit¶
| Key | Description | Default |
|---|---|---|
containerToolkit.enabled |
Whether the Container Toolkit DaemonSet is deployed. Disable if CDI specs and runtime config are managed externally. | true |
containerToolkit.image.* |
Container Toolkit image. | docker.io/rebellions/rbln-container-toolkit:<chart default>pullPolicy: IfNotPresent |
containerToolkit.imagePullSecrets |
Image pull secret(s); required when pulling from a private registry. | [] |
containerToolkit.resources |
CPU/memory requests and limits for the toolkit pod. | {} |
containerToolkit.env |
Environment variables passed to the toolkit, including RBLN_CTK_DAEMON_SOCKET and RBLN_CTK_DAEMON_CONFIG_PATH. |
[] |
containerToolkit.refreshInterval |
How often rbln-ctk-daemon reapplies CDI specs and container runtime configuration on each node. Set to 0s to disable periodic refresh. The value is forwarded to the pod as the RBLN_CTK_DAEMON_REFRESH_INTERVAL environment variable. |
5s |
NPU Feature Discovery¶
| Key | Description | Default |
|---|---|---|
npuFeatureDiscovery.enabled |
Whether the operator deploys NPU Feature Discovery. Disable if you manage NPU node labels through another mechanism. | true |
npuFeatureDiscovery.image.* |
NPU Feature Discovery image. | docker.io/rebellions/rbln-npu-feature-discovery:<chart default>pullPolicy: IfNotPresent |
Metrics Exporter¶
| Key | Description | Default |
|---|---|---|
metricsExporter.enabled |
Whether the Prometheus metrics exporter is deployed. Disable when another telemetry pipeline already covers NPUs. | true |
metricsExporter.image.* |
Metrics exporter image. | docker.io/rebellions/rbln-metrics-exporter:<chart default>pullPolicy: IfNotPresent |
Operator Validator¶
| Key | Description | Default |
|---|---|---|
validator.image.* |
Validator DaemonSet image. | docker.io/rebellions/rbln-npu-operator-validator:<chart default>pullPolicy: IfNotPresent |
validator.imagePullSecrets |
Image pull secret(s); required when pulling from a private registry. | [] |
validator.resources |
CPU/memory requests and limits for the Validator pod. | {} |
validator.env |
Environment variables for the top level Validator process. | [] |
validator.toolkit.env |
Environment variables for the Container Toolkit readiness subcheck. | [] |
validator.driver.env |
Environment variables for the driver readiness subcheck. | [] |
RSMD Node Daemon¶
Each RBLNDriver deploys one RSMD (rbln-smd) DaemonSet. Configure its image through driver.smd.image.* in the Driver Manager table above. RSMD listens on host port 50051. See What the Driver Manager Installs for the deployment and update rules.
Component Logging¶
Each component writes structured logs to stdout. The default log level is info, and the default format is JSON. Configure a component's logging block to override these defaults. The operator maps each configured key to an environment variable; if you omit the block, the image's own defaults apply.
The block is accepted on devicePlugin, draKubeletPlugin, metricsExporter, npuFeatureDiscovery, and sandboxDevicePlugin. Each key maps to the component's own environment variables, which you can also set directly when running the component outside the operator:
| Component | Environment variables |
|---|---|
devicePlugin |
RBLN_DEVICE_PLUGIN_LOG_LEVEL, RBLN_DEVICE_PLUGIN_LOG_FORMAT |
draKubeletPlugin |
RBLN_DRA_DRIVER_LOG_LEVEL, RBLN_DRA_DRIVER_LOG_FORMAT |
metricsExporter |
RBLN_METRICS_EXPORTER_LOG_LEVEL, RBLN_METRICS_EXPORTER_LOG_FORMAT |
npuFeatureDiscovery |
RBLN_NPU_FEATURE_DISCOVERY_LOG_LEVEL, RBLN_NPU_FEATURE_DISCOVERY_LOG_FORMAT |
sandboxDevicePlugin |
RBLN_SANDBOX_DEVICE_PLUGIN_LOG_LEVEL, RBLN_SANDBOX_DEVICE_PLUGIN_LOG_FORMAT |
Do not use trace in production
The metrics exporter and the DRA driver accept trace, which dumps request payloads.
An invalid value falls back to the default and logs a warning rather than failing startup.
The operator's own logging is separate and is configured under operator.logging.*. See Operator Observability.