Inference service configurations
When creating an inference service, you can choose its configuration. The list of available configurations depends on the selected model and its parameters. The configurations automatically include the number and type of graphics processing units (GPU), as well as the amount of vCPU and RAM.
The card for each configuration shows the expected model performance metrics. You can compare the configurations and select the one required for your tasks.
The configuration cannot be changed after the inference service is created.