---
title: "Create an inference service"
sidebar_label: "Create an inference service"
sidebar_position: 5
description: "How to create an inference service"
---

import Formbricks from '@theme/MDXComponents/Formbricks';

# Create an inference service

1. [Select a model](#select-model).
2. [Configure infrastructure](#configure-infrastructure).
3. [Configure the inference service](#configure-inference-service).
4. [Confirm configuration](#confirm-configuration).

## 1. Select a model \{#select-model}

1. In the [control panel](https://my.selectel.ru/ml/default/models-catalog/), in the top menu, click **Products** and select **Dedicated Inference**.

2. Select a [location](/infrastructure/locations.mdx). Once the inference service is created, the location cannot be changed.

3. In the model card, click **Create**.

4. Enter the name of the inference service.

5. To filter inference services in the list, add tags. A tag with the model name is added automatically. To add a new tag, in the **Tags** field, enter the tag and press **Enter**.

6. Optional: enter a description for the inference service. For example, specify its purpose.

7. Click **Continue**.

## 2. Configure infrastructure \{#configure-infrastructure}

1. Specify [model parameters](/dedicated-inference/create/model-parametrs.mdx).

   1.1. Select the data type for the KV cache.

   1.2. Select the maximum context length.

2. Select an [inference service configuration](/dedicated-inference/create/configurations.mdx). When selecting, consider the expected [model performance indicators](/dedicated-inference/create/model-performance-indicators.mdx).

   You cannot change the configuration after creating the inference service.

3. Click **Continue**.

## 3. Configure inference service \{#configure-inference-service}

1. Configure the number of inference instances.

   1.1. To have a fixed number of instances in the service, open the **Fixed** tab and specify the number of instances.

   1.2. To use autoscaling in the service, open the **With autoscaling** tab and set the minimum and maximum number of instances. The number of instances will change automatically only within the specified range depending on the load on the inference service.

   You can [change the number of inference instances](/dedicated-inference/manage/resize-inference-service.mdx#change-number-of-inference-instances) after creating the inference service. For details, see the guide [Scale inference service](/dedicated-inference/manage/resize-inference-service.mdx).

2. Select the disk type for the inference instance.

3. Select an [inference service connection type](/dedicated-inference/create/connection-types.mdx):

   * public — access via a public endpoint from the internet;

   * private — access via a private endpoint only from devices in the inference service's private subnet.

   You cannot change the connection type after creating the inference service.

4. Click **Continue**.

## 4. Confirm configuration \{#confirm-configuration}

1. Review the final inference service configuration.

2. Check the inference service price.

3. Click **Create Inference Service**. Creating an inference service may take about 15 minutes.

<Formbricks />
