
AO
Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek, Guang Lu, Jo\~ao Carvalho
· 1 min read
ResearcharXiv cs.LG
Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing
arXiv:2610.09778v1 Announce Type: cross
Abstract: Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.
Original source
This story was published by arXiv cs.LG and written by Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek, Guang Lu, Jo\~ao Carvalho. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


