---
title: "Dedicated Inference Product Overview"
sidebar_label: "Overview"
sidebar_position: 1
description: "How Dedicated Inference works"
---

import Formbricks from '@theme/MDXComponents/Formbricks';

# Dedicated Inference Product Overview

:::info

The product is in the public testing stage (public preview).

:::

Dedicated Inference is a service for deploying pre-configured ML models with a ready-to-use API in the Selectel infrastructure. [Available models](#available-models) are deployed as isolated inference services with dedicated resources (GPU, vCPU, RAM, disk).

To work with the product, the access control model in Selectel products is supported: [user types](access-control/user-types.mdx), [roles](/access-control/role-reference.mdx), [projects](/access-control/projects/about-projects.mdx) and [project limits and quotas](/access-control/projects/quotas.mdx).

## Use cases \{#tasks-to-be-solved}

* working with models without having to deploy infrastructure yourself;

* selection of turnkey infrastructure for expected or fluctuating workloads. You can compare [model performance indicators](/dedicated-inference/create/model-performance-indicators.mdx) across different [inference service configurations](/dedicated-inference/create/configurations.mdx) and select the required one, as well as set up autoscaling;

* testing and selecting different models for your projects. You can deploy multiple models and compare which one performs better for your tasks;

* integrating models into your own projects via a dedicated endpoint.

## How it works \{#principle-of-operation}

Dedicated Inference uses [Cloud Platform resources](/access-control/projects/about-projects.mdx#selectel-products-that-support-projects). Each model is deployed as an isolated inference service based on a [Managed Kubernetes](/managed-kubernetes/) cluster with GPU. Model weights are stored in [Selectel S3](/s3/).

![](https://423.selcdn.ru/kb/fmc-about-dedicated-inference-architecture-LANG-THEME.png)

For Dedicated Inference fault tolerance, inference services are deployed in clusters in different [locations](/infrastructure/locations.mdx). All clusters with inference services are managed centrally by a separate cluster via [Flux CD](https://fluxcd.io/flux/). Multiple inference services can run in a single cluster — each is isolated from one another and scales independently.

Each inference service uses a suitable inference server depending on the model type. For example, the vLLM inference server is used for large language models (LLMs).

Requests to models are executed via a dedicated endpoint compatible with the OpenAI API.

## How to work with Dedicated Inference \{#manage-product}

You can work with Dedicated Inference in the [control panel](https://my.selectel.ru/ml/default/models-catalog/) or via API. To get started with Dedicated Inference, use the instruction [Dedicated Inference: Quick start](/dedicated-inference/quickstart.mdx).

When creating an inference service, you can select its [configuration](/dedicated-inference/create/configurations.mdx). The list of available configurations depends on the selected model and its [parameters](/dedicated-inference/create/model-parametrs.mdx). When selecting a configuration, you can view expected [model performance indicators](/dedicated-inference/create/model-performance-indicators.mdx).

Also, when creating an inference service, you can select the [inference service connection type](/dedicated-inference/create/connection-types.mdx):

* public — access from the internet;

* private — access from a private network. Access can only be configured via a device in the inference service subnet — for example, [via a cloud server](/dedicated-inference/manage/connection-with-cloud-server-in-private-subnet.mdx).

The endpoint will be automatically generated after the inference service is created and will be available in the Control Panel on the inference service card and page.

Dedicated Inference only supports synchronous mode on the server side. On the client side, you can send synchronous and asynchronous requests. Learn more in the instruction [Inference service interaction modes](/dedicated-inference/manage/interaction-modes.mdx).

Access to an inference service is performed via API keys. API keys are individual for each inference service. You can [manage API keys](/dedicated-inference/manage/manage-api-keys.mdx).

You can scale an inference service depending on the number of requests to the model. To scale an inference service, [change the number of inference instances](/dedicated-inference/manage/resize-inference-service.mdx#change-number-of-inference-instances) — deployed model instances. Learn more in the instruction [Scale inference service](/dedicated-inference/manage/resize-inference-service.mdx).

You can track inference service metrics using the [Metrics](/metrics/about/about-metrics.mdx) service.

## Available models \{#available-models}

You can view the current list of models in the [control panel](https://my.selectel.ru/ml/default/models-catalog/): in the top menu, click **Products** → **Dedicated Inference**.

## Areas of responsibility \{#areas-of-responsibility}

### Selectel provides \{#selectel-provides}

* infrastructure for creating inference services;

* access to models via a public API compatible with OpenAI API;

* the ability to scale inference services;

* inference service monitoring system in the control panel;

* data storage security in compliance with 152-FZ requirements;

* integration with other Selectel services;

* technical support.

### Selectel is not responsible for \{#selectel-is-not-responsible}

* integrating models into your projects;

* the business logic of model operation in your projects.

## Cost \{#price}

Dedicated Inference is billed under the [pay-as-you-go](/balance-and-payments/payment-models.mdx#pay-as-you-go) payment model. Funds are deducted from the balance every hour for the previous hour of using [Cloud Platform resources](/access-control/projects/about-projects.mdx#selectel-products-that-support-projects). The number of tokens does not affect the cost. Learn more in the instruction [Dedicated Inference payment model and pricing](/dedicated-inference/about/payment.mdx).

### What is included in the cost \{#what-included-in-price}

* unlimited number of tokens;

* free domain name for public access to the model.

## Limitations \{#restrictions}

Dedicated Inference does not support:

* working with models in asynchronous mode on the server side;

* uploading custom models to the catalog;

* deploying models from the catalog on-premises, in [Certified Data Center](/certified-data-center-segment/), on [dedicated servers](/dedicated/) and in the [certified cloud](/certified-cloud/).

<Formbricks />
