Skip to main content

General information about the ML Platform product

Selectel ML Platform is a prepared infrastructure for implementing ML development processes: training, deploying ML models, and more. The infrastructure consists of software and hardware components that are configured and ready for use.

All available cloud server configurations are used when selecting ML platform components. Once the platform is connected, you can extend it with your own software components. The following have been tested:

  • ClearML;
  • Kubeflow — for more details on installing Kubeflow, see the Install Kubeflow guide.

There are no additional restrictions on managing the ML Platform cluster from Selectel.

Platform components

By default, the ML Platform consists of:

  • hardware components:
    • Cloud Platform — the base for Managed Kubernetes with NVIDIA® GPUs (Tesla T4, A2, A30, A100, A2000, A5000, GTX 1080⁠, RTX 2080 Ti⁠);
  • software components:
    • Managed Kubernetes clusters with preconfiguration;
    • a domain for accessing the Managed Kubernetes cluster;
    • SSO Keycloak — authorization in internal platform services;
    • Prom Stack — monitoring of platform components;
    • Forecastle — the platform homepage;
    • S3 — storage for datasets and experiment data;
    • Container Registry — container image storage.

In Managed Kubernetes clusters:

  • drivers are installed;
  • nodes are annotated;
  • necessary GPU resources for computing are added;
  • network configured, including Traefik Kubernetes Ingress.

When installing the ClearML platform in a cluster, it is managed directly via the SDK, which is installed in your own IDE. ClearML uses cluster nodes to run ML experiments. The ClearML architecture allows for various component configurations:

  • a single Managed Kubernetes cluster for all ML tasks;
  • several Managed Kubernetes clusters — each for its own task (Inference and Training);
  • connecting a dedicated server as a computational node for ML experiments.

Connect platform

  1. In the control panel, in the top menu, click Products and select ML Platform.
  2. Click Create test request.
  3. Select the data type.
  4. Specify the data volume in GB or MB.
  5. Optional: to help us recommend suitable ways to connect to the ML platform, enter your data source. For example: Selectel, on-premise, or other cloud providers.
  6. Optional: to allow us to take your special data security requirements into account during the test, select the There are additional requirements for ensuring data security in the test checkbox. Describe the requirements in Request comments.
  7. Specify the model size in GB or MB.
  8. Specify the number of people who will use the platform simultaneously.
  9. Select the desired GPU model or select the No GPU model requirements checkbox. GPU specifications can be viewed in the Available GPUs subsection of the Create a cloud server with GPU instructions.
  10. Enter the contact details of a technical specialist. These are required to clarify the technical details of the test.
  11. Optional: enter comments for the application. For example, specify desired tools, components, or requirements for data security during the test.
  12. Click Submit request. A ticket with a request for ML Platform testing will be generated automatically.
  13. Wait for a response from a Selectel employee in the ticket. They will contact you to clarify the details of creating the ML platform.

Cost

The cost of the ML platform is calculated after the application is processed and the configuration is selected. It is formed only from the cost of the platform components: Managed Kubernetes cluster, S3, and Container Registry.