Skip to main content

Foundation Models Catalog: Quick Start

  1. Create an inference service.
  2. Send a test request to the model.

1. Create an inference service

  1. Select a model.
  2. Configure the infrastructure.
  3. Configure the inference service.
  4. Confirm the configuration.

1. Select a model

  1. In the control panel, from the top menu, click Products and select Foundation Models Catalog.

  2. Select a location. Once the inference service is created, the location cannot be changed.

  3. In the model card, click Create.

  4. Enter the name of the inference service.

  5. To filter inference services in the list, add tags. A tag with the model name is added automatically. To add a new tag, in the Tags field, enter the tag and press Enter.

  6. Optional: enter a description for the inference service. For example, specify its purpose.

  7. Click Continue.

2. Configure infrastructure

  1. Set the model parameters.

    1.1. Select the data type for the KV cache.

    1.2. Select the maximum context length.

  2. Select an inference service configuration. When choosing, consider the expected model performance metrics.

    You cannot change the configuration after the inference service is created.

  3. Click Continue.

3. Configure inference service

  1. Configure the number of inference instances.

    1.1. To have a fixed number of instances in the service, open the Fixed tab and specify the number of instances.

    1.2. Чтобы in сервисе использовалось автомасштабирование, откройте вкладку С автомасштабированием and установите минимальное and максимальное количество инстансов. Количество инстансов будет меняться автоматически только in указанном диапазоне in зависимости от нагрузки on inference-сервис.

    Вы можете изменить количество inference-инстансов после создания inference-сервиса. Подробнее in инструкции Масштабировать inference-сервис

  2. Select the disk type for the inference instance.

  3. Select the inference service connection type:

    • public — access via a public endpoint from the internet;

    • private — access via a private endpoint only from cloud platform devices created in the inference service subnet.

    You cannot change the connection type after the inference service is created.

  4. Click Continue.

4. Confirm configuration

  1. Check the final inference service configuration.

  2. Check the inference service price.

  3. Click Create Inference Service. Creating an inference service may take about 15 minutes.

2. Send a test request to the model

The request data structure depends on the model, API type, and request type. Examples of test requests to the model can be copied in the Control Panel: in the top menu, click ProductsInference Services → inference service page → Quick Start tab → in the Test Request block, click .

Подробнее про виды запросов and типы API in инструкции Режимы взаимодействия with inference-сервисом