Dedicated Inference Payment Model and Prices
Balance
To pay for cloud platform resources, depending on the balance type in your account, you can use a unified balance or a cloud platform balance.
You can pay for resources with different types of funds: basic funds or bonuses.
Before paying, top up your balance.
Payment model
The Cloud Platform uses a pay-as-you-go payment model. Every hour, your balance is debited for the previous hour of Cloud Platform resource usage.
Inference service resource billing is generated by project. Each project contains a group of resources: GPU type and quantity.
The resource group cost is updated every astronomical hour.
All Cloud Platform resources for Dedicated Inference are subject to quotas. For quota-based resources, the maximum consumption within an hour is taken into account.
For example, at 13:25, an inference service was created with two inference instances. The configuration of each inference instance: one NVIDIA® A30 GPU with 24 GB of memory. At 13:40, the inference service was scaled — the number of inference instances was reduced to one. The bill for the hour 13:00–14:00 will include the consumption of two NVIDIA® A30 GPUs with 24 GB of memory. For the hour 14:00–15:00, only the consumption of one NVIDIA® A30 GPU with 24 GB of memory will be included, provided the number of inference instances was not increased or other inference services were not created.
If an inference service or the number of inference instances in an inference service is added to a project, the payment for resources will change immediately.
For example, at 13:25, an inference service was created with one inference instance. The configuration of the inference instance: one NVIDIA® A30 GPU with 24 GB of memory. At 13:40, the inference service was scaled — the number of inference instances was increased from one to two. For the hour 13:00–14:00, funds will be deducted for two NVIDIA® A30 GPUs with 24 GB of memory.
Blocking resources if there are insufficient funds on your balance
If there are insufficient funds on your balance at the time of the debit, all cloud platform resources will be automatically blocked — however, you will continue to be charged for them.
To restore access to your resources, you need to top up your balance by the amount of the debt within 14 days after blocking. The debt for resources that were metered during the blocking period will be paid off automatically. Projects are not blocked—you can delete a project in its entirety or its resources via the API.
If you do not top up your balance for the amount of the debt within 14 days after blocking, all Cloud Platform resources will be deleted. Projects will not be deleted in this case.
To ensure you always have enough money on your balance, you can configure balance notifications and automatic balance top-ups.
Prices
The cost depends only on the type and number of GPUs in the inference service configuration. The number of tokens does not affect the cost.
You can view GPU prices when creating an inference service in the control panel.
Accounting documents
After payment, you can receive accounting documents.