SyncAI.news, a Varaisys broadcasting
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
JH

Jitai Hao, Ke Yang, Di Yan, Fan Liu, Qiang Huang, Jun Yu

· 1 min read

ResearcharXiv cs.CV

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

arXiv:2609.02780v2 Announce Type: replace Abstract: Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing full-depth MLLM prefill over incoming frames is prohibitively expensive, incurring substantial computational overhead and causing the KV cache to grow at a rate directly proportional to the prefill depth. To address these challenges, we propose ShallowStream, a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building. During stream processing, ShallowStream maintains an always-on lightweight index using the KV cache of shallow layers. During query-time answering, we leverage the attention scores generated by the shallow layers to score context frames and employ a diversity-aware selection strategy to retrieve precise and comprehensive evidence. ShallowStream achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively. Our code is available at https://github.com/CURRENTF/ShallowStream.

Original source

This story was published by arXiv cs.CV and written by Jitai Hao, Ke Yang, Di Yan, Fan Liu, Qiang Huang, Jun Yu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News