Skip to main content

Foundation Models Catalog: Quick Start

  1. Create an inference service.
  2. Send a test request to the model.

1. Create an inference service

  1. Select a model.
  2. Configure the infrastructure.
  3. Configure the inference service.
  4. Confirm the configuration.

1. Select a model

  1. In the control panel, from the top menu, click Products and select Foundation Models Catalog.

  2. In the model card, click Create.

  3. Enter the name of the inference service.

  4. To filter inference services in the list, add tags. A tag with the model name is added automatically. To add a new tag, in the Tags field, enter the tag and press Enter.

  5. Optional: enter a description for the inference service. For example, specify its purpose.

  6. Click Continue.

2. Configure the infrastructure

  1. Set the model parameters.

    1.1. Select the data type for the KV cache.

    1.2. Select the maximum context length.

  2. Select an inference service configuration. When choosing, consider the expected model performance metrics.

    Once the inference service is created, the configuration cannot be changed.

  3. Click Continue.

3. Set up the inference service

  1. Configure the number of inference instances.

    1.1. To have a fixed number of instances in the service, open the Fixed tab and specify the number of instances.

    1.2. To use autoscaling in the service, open the Auto-scaling tab and set the minimum and maximum number of instances. The number of instances will change automatically only within the specified range depending on the inference service load.

    You can change the number of inference instances after creating the inference service. For details, see the Scaling an inference service instruction.

  2. Select the volume type for the inference instance.

  3. Click Continue.

4. Confirm the configuration

  1. Check the final inference service configuration.

  2. Check the price of the inference service.

  3. Click Create Inference Service. Creating an inference service may take about 15 minutes.

2. Send a test request to the model

The request data structure depends on the model, the API type, and the type of request. You can copy test request examples for the model in the Control panel: in the top menu, click ProductsInference Services → the inference service page → the Quick start tab → in the Test request block, click .

Learn more about request types and API types in the Interaction modes with the inference service instruction.