Skip to main content

Graphics Processing Units (GPU)

You can add GPUs (graphics processing units) to a cloud server when creating a server or to an existing server.

GPUs are used as dedicated PCI devices inside a cloud server.

GPUs are available in fixed and custom configurations of the GPU line.

GPU lineup configurations and custom configurations with GPUs can be used with a local or network boot volume.

For cloud servers with a local disk, you can use:

  • NVIDIA® A100 40Gb;
  • NVIDIA® A100 80Gb;
  • NVIDIA® A30 24Gb;
  • NVIDIA® RTX 4090 48Gb;
  • NVIDIA® A5000 24Gb;
  • NVIDIA® RTX 6000 Ada;
  • NVIDIA® PRO 6000 96Gb;
  • NVIDIA® H200 141Gb.

If you need a server with a pre-configured set of tools and libraries for machine learning and data analysis, use the AI Marketplace.

Available GPUs

MemoryCUDA CoresTensor Cores

NVIDIA® A100 40Gb

NVIDIA® A100 40Gb NVLink (upon request)

40 GB
HBM2

6,192432
NVIDIA® A100 80Gb80 GB
HBM2
6,912432
NVIDIA® Tesla T4 16Gb16 GB
GDDR6
2,560320
NVIDIA® A30 24Gb24 GB
HBM2
3,804224
NVIDIA® A2 16Gb
(updated equivalent of
NVIDIA® Tesla T4)
16 GB
GDDR6
1,28040
NVIDIA® GTX 1080 8Gb8 GB
GDDR5X
2,560✗
NVIDIA® RTX 2080 Ti 11Gb11 GB
GDDR6
4,352544
NVIDIA® RTX 4090 24Gb24 GB
GDDR6X
16,384512
NVIDIA® RTX 4090 48Gb48 GB
GDDR6X
16,384512
NVIDIA® RTX 6000 Ada 48Gb
(equivalent to L40)
48 GB
GDDR6X
18,176568
NVIDIA® A2000 6Gb
(equivalent to RTX 3060)
6 GB
GDDR6
3,328104
NVIDIA® A5000 24Gb
(equivalent to RTX 3080)
24 GB
GDDR6
8,192256
NVIDIA® H100 80Gb80 GB
HBM3
16,896528
NVIDIA® H200 141Gb141 GB
HBM3e
16,896528
NVIDIA® L4 24Gb24 GB
GDDR6
20,480640
NVIDIA® PRO 6000 96Gb48 GB
GDDR7
18,432576
NVIDIA® A4000 16Gb16 GB
GDDR6
6,144192

You can view the current list of GPUs in the control panel: in the top menu, click Products → Cloud servers → click Create server.

You can check GPU availability by location in the GPU section of the Cloud Server Availability by Location instructions.

NVIDIA® A100 40Gb

Features maximum performance for AI, HPC, and data processing. Suitable for deep learning, scientific research, and data analytics.

Based on the Ampere® architecture, with 40 GB HBM2 memory and bandwidth up to 1.5 GB/s. Learn more about the NVIDIA® A100 40Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 40 GB are available, with 6 to 48 vCPUs and 87 to 700 GB RAM.

In custom configurations — from 1 to 8 GPUs × 40 GB, with vCPUs from 2 to 32, RAM from 512 MB to 256 GB.

For NVIDIA® A100 40Gb GPUs connected using NVLink technology.

NVLink accelerates data transfer speeds when combining GPUs compared to the PCIe interface. GPUs interconnected with NVLink allow using more memory and increase server performance for complex computations, such as training large language ML models.

NVLink works with NVIDIA® A100 40Gb — GPUs based on the Ampere® architecture, with 40 GB HBM2 memory and bandwidth up to 1.5 GB/s. Learn more about the NVIDIA® A100 40Gb specifications and read the description of NVLink technology in the NVIDIA® documentation.

NVIDIA® A100 40Gb NVLink are available upon request — submit a ticket.

NVIDIA® A100 80Gb

Features maximum performance for AI, HPC, and data processing, as well as large memory capacity for resource-intensive tasks. Suitable for deep learning, scientific research, and data analytics.

Based on the Ampere® architecture, with 80 GB HBM2 memory and bandwidth up to 1.5 GB/s. Learn more about the NVIDIA® A100 80Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 80 GB are available, with 12 to 96 vCPUs, 128 GB to 1 TB RAM, and local disk from 128 GB to 6.88 TB.

In custom configurations — from 1 to 8 GPUs × 80 GB, with 12 to 192 vCPUs, 64 GB to 1 TB RAM, and local disk from 256 GB to 3.36 TB.

NVIDIA® Tesla T4 16Gb

Suitable for Machine Learning and Deep Learning, inference, graphics processing, and video rendering. Works with most AI frameworks and is compatible with all types of neural networks.

Based on the Turing® architecture, with 16 GB GDDR6 memory and bandwidth up to 300 GB/s. Learn more about the NVIDIA® Tesla T4 16Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 16 GB are available, with 4 to 24 vCPUs and 32 to 320 GB RAM.

In custom configurations — from 1 to 4 GPUs × 16 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® A30 24Gb

Suitable for AI inference, HPC, language processing, conversational AI, and recommendation systems.

Based on the Ampere® architecture, with 24 GB HBM2 memory and bandwidth up to 933 GB/s. Learn more about the NVIDIA® A30 24Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 2 GPUs × 24 GB are available, with 16 to 48 vCPUs and 64 to 320 GB RAM.

In custom configurations — from 1 to 2 GPUs × 24 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® A2 16Gb

Entry-level GPU. Suitable for basic inference, video and graphics, Edge AI (edge computing), Edge video, and mobile cloud gaming.

Based on the Ampere® architecture, with 16 GB GDDR6 memory and bandwidth up to 200 GB/s. Learn more about the NVIDIA® A2 16Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 16 GB are available, with 12 to 48 vCPUs and 32 to 320 GB RAM.

In custom configurations — from 1 to 4 GPUs × 16 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® GTX 1080 8Gb

A powerful and energy-efficient GPU. The solution is built with FinFET technology and GDDR5X memory. Dynamic load balancing helps distribute tasks so resources don't sit idle. Features maximum performance for display output, VR, ultra-high resolution settings, and data processing.

Based on the Pascal® architecture, with 8 GB GDDR5X memory and bandwidth up to 320 GB/s. Learn more about the NVIDIA® GTX 1080 8Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 8 GB are available, with 8 to 28 vCPUs and 24 to 96 GB RAM.

In custom configurations — from 1 to 8 GPUs × 8 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® RTX 2080 Ti 11Gb

High-performance GPU for complex graphics tasks. Suitable for:

  • high-resolution video processing;
  • creating 3D models;
  • rendering and photo editing;
  • neural network training;
  • performing complex computations in the field of artificial intelligence;
  • processing large volumes of data.

Based on the Turing® architecture, with 11 GB GDDR6 memory and bandwidth up to 616 GB/s. Learn more about the NVIDIA® RTX 2080 Ti 11Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 11 GB are available, with 2 to 48 vCPUs and 32 to 320 GB RAM.

In custom configurations — from 1 to 4 GPUs × 11 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® RTX 4090 24Gb

A high-performance GeForce GPU suitable for:

  • professional design and 3D modeling;
  • video processing and rendering;
  • ML tasks (model training and inference);
  • working with large language models (LLMs);
  • scientific and engineering computations (for example, in climate modeling or bioinformatics).

Based on the Ada Lovelace® architecture, with 24 GB GDDR6X memory and bandwidth up to 1008 GB/s. Learn more about the NVIDIA® RTX 4090 specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 24 GB are available, with 4 to 64 vCPUs and 16 to 356 GB RAM.

In custom configurations — from 1 to 4 GPUs × 24 GB, with 2 to 32 vCPUs and 4 to 256 GB RAM.

NVIDIA® RTX 4090 48Gb

A high-performance GeForce GPU with increased memory capacity compared to the NVIDIA® RTX 4090 24 Gb, suitable for:

  • professional design and 3D modeling;
  • video processing and rendering;
  • ML tasks (model training and inference);
  • working with large language models (LLMs);
  • scientific and engineering computations (for example, in climate modeling or bioinformatics).

Based on the Ada Lovelace® architecture, with 48 GB GDDR6X memory and bandwidth up to 1008 GB/s. Learn more about the NVIDIA® RTX 4090 specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 48 GB are available, with 12 to 192 vCPUs, 64 to 896 GB RAM, and local disk from 64 to 800 GB.

In custom configurations — from 1 to 8 GPUs × 48 GB, with 12 to 192 vCPUs, 64 to 896 GB RAM, and local disk from 50 to 800 GB.

NVIDIA® RTX 6000 Ada 48Gb

Professional GPU for compute and graphics power. Suitable for ML tasks, rendering, scientific computing, and high-performance visualization.

Based on the Ada Lovelace® architecture, with 48 GB GDDR6X memory and bandwidth up to 960 GB/s. Learn more about the NVIDIA® RTX 6000 Ada specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 48 GB are available, with 12 to 96 vCPUs, 64 to 450 GB RAM, and local disk from 64 GB to 2 TB.

In custom configurations — from 1 to 4 GPUs × 48 GB, with 12 to 96 vCPUs, 64 to 450 GB RAM, and local disk from 64 GB to 3.52 TB.

NVIDIA® A2000 6Gb

Energy-efficient GPU for compact workstations. Suitable for AI, graphics, and video rendering.

Based on the Ampere® architecture, with 6 GB GDDR6 memory and bandwidth up to 288 GB/s. Learn more about the NVIDIA® A2000 6Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 4 GPUs × 6 GB are available, with 6 to 24 vCPUs and 16 to 320 GB RAM.

In custom configurations — from 1 to 4 GPUs × 6 GB, with 2 to 32 vCPUs and 512 MB to 256 GB RAM.

NVIDIA® A5000 24Gb

Versatile GPU suitable for any tasks within its performance range.

Based on the Ampere® architecture, with 24 GB GDDR6 memory and bandwidth up to 768 GB/s. Learn more about the NVIDIA® A5000 24Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 24 GB are available, with 8 to 128 vCPUs, 16 to 700 GB RAM, and local disk from 64 GB to 1 TB.

In custom configurations — from 1 to 8 GPUs × 24 GB, with 2 to 128 vCPUs, 2 to 731 GB RAM, and local disk from 20 GB to 2 TB.

NVIDIA® H100 80Gb

A powerful GPU suitable for AI, HPC, and scalable computing.

Based on the Hopper™ architecture, with 80 GB HBM3 memory and bandwidth up to 3 TB/s. Learn more about the NVIDIA® H100 80Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 2 GPUs × 80 GB are available, with 12 to 48 vCPUs and 128 to 256 GB RAM.

In custom configurations — from 1 to 2 GPUs × 80 GB, with 2 to 48 vCPUs and 2 GB to 256 GB RAM.

NVIDIA® H200 141GB

A professional GPU suitable for:

  • accelerating generative AI;
  • high-performance computing (HPC);
  • large language model (LLM) inference;
  • model fine-tuning;
  • image and video generation.

Based on the Hopper™ architecture, with 141 GB HBM3 memory and bandwidth up to 4.8 TB/s. Learn more about the NVIDIA® H200 141Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed and custom configurations of the GPU line, 1 to 8 GPUs × 141 GB are available, with 12 to 192 vCPUs, 120 GB to 1 TB of RAM, and a local drive from 256 GB to 3 TB.

Two NVIDIA® H200 141Gb GPUs connected using NVLink technology.

NVLink accelerates data transfer speeds when combining GPUs compared to the PCIe interface. GPUs interconnected with NVLink allow using more memory and increase server performance for complex computations, such as training large language models (LLMs).

NVLink works with NVIDIA® H200 — GPUs based on the Hopper™ architecture, with 141 GB HBM3 memory and bandwidth up to 4.8 TB/s. Learn more about the NVIDIA® H200 141Gb specifications in the NVIDIA® documentation.

In fixed configurations of the GPU line, 2 GPUs × 141 GB are available, with 24 to 32 vCPUs, 240 GB RAM, and a 256 GB local drive. Cloud servers with NVIDIA® H200 NVLink also run on dedicated cores and support 10 Gbps fast networks

In fixed configurations of the Dedicated line, 4 GPUs × 141 GB are available, with 48 to 96 vCPUs, 480 to 720 GB RAM, and a 512 GB to 1 TB local drive.

NVIDIA® L4 24Gb

A versatile GPU for accelerating AI/ML workloads, video processing, streaming, and VDI. Suitable for running modern large language models (LLMs) and multimodal models.

Based on the Ada Lovelace® architecture, with 24 GB GDDR6 memory and bandwidth up to 3 TB/s. Learn more about the NVIDIA® L4 24Gb specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed configurations of the GPU line, 1 to 8 GPUs × 24 GB are available, with 8 to 128 vCPUs and 32 GB to 512 GB RAM.

In custom configurations — from 1 to 8 GPUs × 24 GB, with 8 to 256 vCPUs and 64 GB to 640 GB RAM.

NVIDIA® PRO 6000 96Gb

Professional GPU:

  • accelerating generative AI;
  • inference of large language models (LLMs);
  • fine-tuning models;
  • generating images and videos;
  • 3D rendering and video processing.

Based on the Blackwell® architecture, with 96 GB GDDR7 memory and bandwidth up to 1.6 TB/s. Learn more about the NVIDIA® RTX 6000 Pro specifications in the NVIDIA® documentation. When using multiple GPUs, data transfer between them is carried out via PCIe.

In fixed and custom configurations of the GPU line, 1 to 8 GPUs × 96 GB are available, with 16 to 256 vCPUs, 120 GB to 1 TB RAM, and local disk from 256 GB to 3 TB.

NVIDIA® A4000 16Gb

A versatile GPU suitable for efficient work with graphics, computations, and machine learning tasks.

Based on the Ampere® architecture, with 16 GB GDDR6 memory and bandwidth up to 448 GB/s. Learn more about the NVIDIA® RTX A4000 16Gb specifications in the NVIDIA® documentation.

In fixed and custom configurations of the GPU line, 1 to 8 GPUs × 16 GB are available, with 4 to 64 vCPUs and 16 to 256 GB RAM.

Create a cloud server with GPU

Follow the Create a cloud server instructions.

When creating a server, select:

  • source — ready-made GPU-optimized images marked as GPU optimized in the version list. The images contain the drivers required to work with GPUs. If you select a different source, for stable NVIDIA® GPU operation, you will need to install drivers on the server yourself;
  • configuration — a fixed or custom configuration from the GPU line with at least 2 vCPUs.