
KT
Kaizhen Tan, Yang Feng, Heqing Du, Hanzhe Hong, Siru Tao, Xin Xu
· 1 min read
ResearcharXiv cs.CV
Confidence under Visual Token Pruning: Removed Evidence and Risk-Controlled Token Budgets for MLLMs
arXiv:2604.12035v4 Announce Type: replace
Abstract: Visual token pruning speeds up multimodal large language models (MLLMs) by keeping a small subset of the visual tokens, and pruning methods are compared by the accuracy they retain. We study what pruning does to the confidence of these models, across common selectors, several MLLMs, and different output formats. Pruning errors concentrate on questions whose evidence the selector removed, and confidence does not register the removal. When the queried object loses all of its tokens, accuracy on these questions drops from 59% to 17%, while confidence stays at the unpruned level. Returning a few object tokens to the kept set recovers most of the lost accuracy. Selectors that keep the most attended tokens remove such evidence most often and produce confident errors, which temperature scaling cannot re-rank. Selectors that avoid keeping redundant tokens stay close to the calibration of the unpruned model. We then use the confidence of the pruned model to set a per-question token budget. The model answers with few tokens first and again with all tokens when its confidence is low. Conformal risk control sets the threshold to bound the expected deviation from the unpruned model. With coverage-based selection, this cascade needs about a third of the prefill tokens of the unpruned model, while with FastV it needs more than the unpruned model. The savings come mainly from how well the confidence ranks the answers that differ from the unpruned ones.
Original source
This story was published by arXiv cs.CV and written by Kaizhen Tan, Yang Feng, Heqing Du, Hanzhe Hong, Siru Tao, Xin Xu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


