Inference service configurations
When creating an inference service, you can select its configuration. The list of available configurations depends on the selected model and its parameters. The configurations automatically include the appropriate number and type of graphics processing units (GPU), amount of vCPU, and RAM.
The card for each configuration shows the expected model performance metrics. You can compare configurations and choose the one required for your tasks.
The configuration cannot be changed after the inference service is created.