InferenceBench

llm.inference.sharegpt-v3

1 entry. Pareto frontier computed on throughput_tok_per_s (higher is better) vs. ttft_p50_ms (lower is better). Rows marked P are on the frontier.

1 of 1 matching
Model Engine Hardware Quant TTFT P50 (ms) TTFT P99 (ms) Throughput (tok/s) $/M tokens J/token Power avg (W) Power peak (W) WER mean J / audio s Envelope
P Qwen/Qwen2.5-72B-Instruct vllm 0.22.1 8x NVIDIA H100 80GB HBM3 bf16 45.95 55.79 125 15.92 1,995 2,200 JSON