Skip to main content

Inference Service model performance metrics

Use metrics to compare model performance across different configurations and select the necessary one. Metrics are displayed in the card for each configuration when creating an inference service.

Avg Time to First Token

Average time from request receipt to generation of the first token in milliseconds

Avg Request Throughput

Average number of requests processed per second

Output Token Throughput

Average number of generated tokens per second

Request Latency

Average time from request receipt to full response completion in seconds