Inference service configurations
When creating an inference service, you can select its configuration. The list of available configurations depends on the selected model and its parameters. The configurations automatically include the appropriate number and type of graphics processing units (GPU), number of vCPU, and RAM.
Each configuration card lists the expected model performance indicators. You can compare configurations and select the one you need for your tasks.
The configuration cannot be changed after the inference service is created.