Skip to main content

Inference service model performance indicators

You can compare model performance indicators across different configurations and choose the one you need. Performance indicators are displayed in the card of each configuration when creating an inference service.

Avg Time to First Token

Average time from receiving a request to generating the first token in milliseconds

Avg Request Throughput

Average number of processed requests per second

Output Token Throughput

Average number of generated tokens per second

Request Latency

Average time from receiving a request to full response completion in seconds