SyncAI.news, a Varaisys broadcasting
OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing
LQ

Long Qian, Bingke Zhu, Jiaqi Wei, Yingying Chen, Jinqiao Wang

· 1 min read

ResearcharXiv cs.CV

OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing

arXiv:2609.32780v1 Announce Type: new Abstract: Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be exposed to language computation for a given question? We introduce OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector. For each image-prompt pair, the projector selects a sparse set of intermediate visual depths and reuses the resulting global preference to guide both local patch fusion and dynamic visual injection into the language model. This design enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks. Across eight image-based benchmarks, OmniMoE-VL achieves an average score of 85.9 with 28B total and 9B activated parameters. Controlled comparisons show that the routed visual interface provides the dominant architectural gain, while matched route and component controls, same-image route analysis, and route interventions further support the value of coupling and question-conditioned visual access.

Original source

This story was published by arXiv cs.CV and written by Long Qian, Bingke Zhu, Jiaqi Wei, Yingying Chen, Jinqiao Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News