Skip to main content

Drivers for GPU node groups in a Managed Kubernetes cluster

You can create Managed Kubernetes clusters on a cloud server with GPU:

Kernel modules in pre-installed drivers

Kernel modules in pre-installed drivers can be:

  • proprietary — used in GPUs with Pascal, Maxwell, and Volta architecture;
  • open-source — used in GPUs with Turing architecture and all subsequent architecture generations.

Among the available GPUs, only the NVIDIA® GTX 1080 GPU with Pascal-based architecture has proprietary kernel modules, while the rest have open-source ones.

For more information about kernel modules, see the Kernel Modules section in the NVIDIA® documentation.

Install drivers

To install the driver yourself, use NVIDIA® GPU Operator.

  1. Connect to the cluster.

  2. Install the Helm package manager version 3.7.0 or higher.

  3. Add the nvidia repository to Helm:

    helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
  4. Update the nvidia repository in Helm:

    helm repo update
  5. Install NVIDIA GPU Operator and specify the required GPU driver version:

    helm install \
    --namespace gpu-operator \
    --create-namespace \
    --set driver.version=<driver_version> \
    gpu-operator nvidia/gpu-operator

    Specify <driver_version> — NVIDIA® driver version. You can find it in the NVIDIA GPU Driver row of the GPU Operator Component Matrix table in the NVIDIA® documentation.

  6. To verify that NVIDIA GPU Operator and the GPU driver are installed correctly, run a GPU application. For example, the CUDA VectorAdd vector addition application:

    cat << EOF | kubectl create -f -
    apiVersion: v1
    kind: Pod
    metadata:
    name: cuda-vectoradd
    spec:
    restartPolicy: OnFailure
    containers:
    - name: cuda-vectoradd
    image: "nvidia/samples:vectoradd-cuda11.2.1"
    resources:
    limits:
    nvidia.com/gpu: 1
    EOF
  7. Make sure that the CUDA VectorAdd application has completed successfully — the pod status should be Completed:

    kubectl get pods

    In the response, the cuda-vectoradd pod will have the status Completed:

    NAME READY STATUS RESTARTS AGE
    cuda-vectoradd 0/1 Completed 0 51s