Skip to main content

Create a Managed Kubernetes cluster on a cloud server with a GPU

You can add GPUs (graphics processing units) to a Managed Kubernetes cluster on a cloud server when creating a Managed Kubernetes cluster on a cloud server or adding a node group on a cloud server.

You can check GPU availability in locations in the GPU for Managed Kubernetes subsection of the Product availability by location instructions.

On nodes with GPUs, you can use pre-installed drivers or install drivers manually. For GPU node groups without drivers, cluster autoscaling is not available.

Create a cluster on a cloud server with a GPU

Follow the Create a Managed Kubernetes cluster on a cloud server instructions.

Configure:

  • configuration — a fixed configuration of a node group with a GPU;
  • GPU drivers — by default, the GPU Drivers toggle is enabled and the cluster uses pre-installed drivers. To install GPU drivers manually, disable the GPU Drivers toggle.

Available GPUs

MemoryCUDA coresTensor cores

NVIDIA® A100 40Gb

40 GB
HBM2

6192432
NVIDIA® A100 80Gb80 GB
HBM2
6912432
NVIDIA® Tesla T4 16Gb16 GB
GDDR6
2560320
NVIDIA® A30 24Gb24 GB
HBM2
3804224
NVIDIA® A2 16Gb
(updated equivalent of
NVIDIA® Tesla T4)
16 GB
GDDR6
128040
NVIDIA® GTX 1080 8Gb8 GB
GDDR5X
2560✗
NVIDIA® RTX 2080 Ti 11Gb11 GB
GDDR6
4352544
NVIDIA® RTX 4090 24Gb24 GB
GDDR6X
16384512
NVIDIA® RTX 4090 48Gb48 GB
GDDR6X
16384512
NVIDIA® RTX 6000 Ada
(equivalent to L40)
48 GB
GDDR6X
18176568
NVIDIA® A2000 6Gb
(equivalent to RTX 3060)
6 GB
GDDR6
3328104
NVIDIA® A5000 24Gb
(equivalent to RTX 3080)
24 GB
GDDR6
8192256
NVIDIA® H100 80Gb80 GB
HBM3
16896528
NVIDIA® H200 141Gb141 GB
HBM3e
16896528
NVIDIA® L4 24Gb24 GB
GDDR6
20480640
NVIDIA® PRO 6000 96Gb48 GB
GDDR7
18432576
NVIDIA® A4000 16Gb16 GB
GDDR6
6,144192

You can view the current list of GPUs in the Control panel: in the top menu, click Products → Managed Kubernetes → Create cluster → the Node groups step → Cloud server Node configuration → Fixed with GPU.

You can check GPU availability in locations in the GPU for Managed Kubernetes subsection of the Product availability by location instructions.

NVIDIA® A100 40Gb

Delivers maximum performance for AI, HPC, and data processing. Suitable for deep learning, scientific research, and data analytics.

Based on the Ampere® architecture, bandwidth up to 1.5 GB/s. For more information about the specifications of NVIDIA® A100 40Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 40 GB are available, with 6 to 48 vCPUs and 87 to 704 GB RAM.

NVIDIA® A100 80Gb

Delivers maximum performance for AI, HPC, and data processing, as well as a large memory capacity for resource-intensive tasks. Suitable for deep learning, scientific research, and data analytics.

Based on the Ampere® architecture, with 80 GB HBM2 memory and bandwidth up to 1.5 GB/s. For more information about the specifications of NVIDIA® A100 80Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 80 GB are available, with 12 to 192 vCPUs and 128 GB to 1 TB RAM.

NVIDIA® Tesla T4 16Gb

Suitable for Machine Learning and Deep Learning, inference, graphics processing, and video rendering. Compatible with most AI frameworks and all types of neural networks.

Based on the Turing® architecture, bandwidth up to 300 GB/s. For more information about the specifications of NVIDIA® Tesla T4 16Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 16 GB are available, with 4 to 24 vCPUs and 32 to 320 GB RAM.

NVIDIA® A30 24Gb

Suitable for AI inference, HPC, natural language processing, conversational AI, and recommendation systems.

Based on the Ampere® architecture, bandwidth up to 933 GB/s. For more information about the specifications of NVIDIA® A30 24Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 2 GPUs × 24 GB are available, with 16 to 48 vCPUs and 64 to 320 GB RAM.

NVIDIA® A2 16Gb

An entry-level GPU. Suitable for basic inference, video and graphics, Edge AI (edge computing), Edge video, and mobile cloud gaming.

Based on the Ampere® architecture, bandwidth up to 200 GB/s. For more information about the specifications of NVIDIA® A2 16Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 16 GB are available, with 12 to 48 vCPUs and 32 to 320 GB RAM.

NVIDIA® GTX 1080 8Gb

A high-performance and energy-efficient GPU. Built using FinFET technology and GDDR5X memory. Dynamic load balancing helps distribute tasks so resources do not sit idle. Delivers maximum performance for display output, VR, ultra-high resolution settings, and data processing.

Based on the Pascal® architecture, bandwidth up to 320 GB/s. For more information about the specifications of NVIDIA® GTX 1080 8Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 8 GB are available, with 8 to 28 vCPUs and 24 to 96 GB RAM.

NVIDIA® RTX 2080 Ti 11Gb

A high-performance GPU for demanding graphics tasks. Suitable for:

  • high-resolution video processing;
  • 3D modeling;
  • rendering and photo processing;
  • training neural networks;
  • complex artificial intelligence computing;
  • processing large volumes of data.

Based on the Turing® architecture, bandwidth up to 616 GB/s. For more information about the specifications of NVIDIA® RTX 2080 Ti 11Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 11 GB are available, with 2 to 48 vCPUs and 32 to 320 GB RAM.

NVIDIA® RTX 4090 24Gb

A high-performance GeForce series GPU. Suitable for professional design and 3D modeling, video editing, rendering, ML tasks (model training and inference), working with language models (LLMs), and scientific and engineering computing (for example, in climate modeling or bioinformatics).

Based on the Ada Lovelace® architecture, bandwidth up to 1008 GB/s. For more information about the specifications of NVIDIA® RTX 4090, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 24 GB are available, with 4 to 64 vCPUs and 16 to 356 GB RAM.

NVIDIA® RTX 4090 48Gb

A high-performance GeForce series GPU with increased memory compared to the NVIDIA® RTX 4090 24 Gb, suitable for:

  • professional design and 3D modeling;
  • video editing and rendering;
  • ML tasks (model training and inference);
  • working with language models (LLMs);
  • scientific and engineering computing (for example, in climate modeling or bioinformatics).

Based on the Ada Lovelace® architecture, with 48 GB GDDR6X memory and bandwidth up to 1008 GB/s. For more information about the specifications of NVIDIA® RTX 4090, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 48 GB are available, with 12 to 192 vCPUs, 64 to 896 GB RAM, and a 64 to 800 GB local drive.

NVIDIA® RTX 6000 Ada 48Gb

A professional GPU for computing and graphics power. Suitable for ML tasks, rendering, scientific computing, and high-performance visualization.

Based on the Ada Lovelace® architecture, with 48 GB GDDR6X memory and bandwidth up to 960 GB/s. For more information about the specifications of NVIDIA® RTX 6000 Ada, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 48 GB are available, with 12 to 96 vCPUs and 64 to 450 GB RAM.

NVIDIA® A2000 6Gb

An energy-efficient GPU for compact workstations. Suitable for AI, graphics, and video rendering.

Based on the Ampere® architecture, bandwidth up to 288 GB/s. For more information about the specifications of NVIDIA® A2000 6Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 4 GPUs × 6 GB are available, with 6 to 24 vCPUs and 16 to 320 GB RAM.

NVIDIA® A5000 24Gb

A versatile GPU suitable for any task within its performance capacity.

Based on the Ampere® architecture, bandwidth up to 768 GB/s. For more information about the specifications of NVIDIA® A5000 24Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 24 GB are available, with 8 to 128 vCPUs and 16 to 700 GB RAM.

NVIDIA® H100 80Gb

A powerful GPU suitable for AI, HPC, and scalable computing.

Based on the Hopper™ architecture, with 80 GB HBM3 memory and bandwidth up to 3 TB/s. For more information about the specifications of NVIDIA® H100 80Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 2 GPUs × 80 GB are available, with 12 to 48 vCPUs and 128 to 256 GB RAM.

NVIDIA® H200 141Gb

A professional GPU for:

  • accelerating generative AI;
  • high-performance computing (HPC);
  • large language model (LLM) inference;
  • fine-tuning models;
  • image and video generation.

Based on the Hopper™ architecture, with 141 GB HBM3 memory and bandwidth up to 4.8 TB/s. For more information about the specifications of NVIDIA® H200 141Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 141 GB are available, with 12 to 192 vCPUs and 120 GB to 1 TB RAM.

NVIDIA® L4 24Gb

A universal GPU for accelerating AI/ML workloads, video processing, streaming, and VDI. Suitable for running modern language models (LLMs) and multimodal models.

Based on the Ada Lovelace® architecture, with 24 GB GDDR6 memory and bandwidth up to 3 TB/s. For more information about the specifications of NVIDIA® L4 24Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 24 GB are available, with 8 to 128 vCPUs and 32 GB to 512 GB RAM.

NVIDIA® PRO 6000 96Gb

A professional GPU for:

  • accelerating generative AI;
  • language model (LLM) inference;
  • fine-tuning models;
  • image and video generation;
  • 3D rendering and video processing.

Based on the Blackwell® architecture, with 96 GB GDDR7 memory and bandwidth up to 1.6 TB/s. For more information about the specifications of NVIDIA® RTX 6000 Pro, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 96 GB are available, with 16 to 256 vCPUs and 120 GB to 1 TB RAM.

NVIDIA® A4000 16Gb

A versatile GPU suitable for efficient work with graphics, compute, and machine learning tasks.

Based on the Ampere® architecture, with 16 GB GDDR6 memory and bandwidth up to 448 GB/s. For more information about the specifications of NVIDIA® RTX A4000 16Gb, see the NVIDIA® documentation.

In fixed Managed Kubernetes cluster configurations, 1 to 8 GPUs × 16 GB are available, with 4 to 64 vCPUs and 16 to 256 GB RAM.