Skip to main content

Release notes for Dedicated Inference

2026

August
  • added the ability to track inference service metrics using the Metrics service;

  • added new models:

    You can view the current list of available models in the Control Panel: in the top menu, click Products → Dedicated Inference;

  • simplified the inference service cost calculation: it is now based only on GPU consumption. Previously, the cost was composed of several resources: GPU, vCPU, RAM, and disk. For more details, see the Dedicated Inference Payment Model and Pricing instructions.

July
  • added the ability to sort configurations when creating an inference service. Now, at the configuration selection step, they can be sorted in ascending and descending order of price as well as performance;

  • Dedicated Inference is now available in the ru-6 pool. Also added the ability to select a location in the model catalog when creating an inference service;

  • added the ability to connect to an inference service from a private network. For more details, see the Inference service connection types instructions.

June
April
  • moved Dedicated Inference from the private preview stage to the public preview stage.