SyncAI.news, a Varaisys broadcasting
Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity
VA

Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu

· 1 min read

ResearcharXiv cs.AI

Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

arXiv:2606.22790v4 Announce Type: replace-cross Abstract: Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along \emph{six} dimensions: model size, temporal resolution, encoder token stride, low-rank adaptation capacity, weight precision and sparsity pattern. All axes are jointly optimized using NSGA-III with respect to three deployment objectives: word error rate (WER), inference FLOPs, and memory footprint. Across 50 of the 1,680 candidate configurations evaluated, we characterize the conditional effect of each axis and identify compression combinations that dominate naive single-axis scaling, while finding that 1:4 structured sparsity fails to recover acceptable accuracy under the tested recovery budgets. We report measured WER and resident memory, use analytical EffFLOPs as the search-time compute surrogate, and separately validate representative inference configurations using measured real-time factor (RTF).

Original source

This story was published by arXiv cs.AI and written by Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News