
AR
Arefeh Rezaei
· 1 min read
ResearcharXiv cs.CV
Video Captioning in Low-Light Conditions through Efficient Uncertainty-Aware Caption Correction
arXiv:2609.31697v1 Announce Type: new
Abstract: Low-light conditions can significantly degrade the ability of vision-language models (VLMs) to accurately describe human actions in videos. In this work, I propose an efficient uncertainty-aware representation correction framework for improving captions generated by VideoChat2 under real-world low-light conditions. Instead of fine-tuning the underlying VLM, the proposed framework introduces a lightweight sparse Gaussian process-based error estimation module between the projection layer and the language model to correct the intermediate representation. The correction module learns to estimate the residual between the original projected representation and a verified target representation, which is then adaptively scaled using a newly formulated uncertainty-aware coefficient and added to the original representation. To further improve residual estimation, I introduce a partitioned combined-kernel design. The correction model is trained separately using only 44 samples from the ARID dataset and requires only a small additional computational overhead during inference. The effectiveness of the proposed correction is evaluated through quantitative residual prediction and qualitative analysis of the generated captions. Although VideoChat2 is used in my experiments, the proposed framework is designed to be applicable to other compatible VLM architectures. \textbf{Code Availability: The implementation accompanying this work is publicly available at}:\href{https://github.com/areferezaee/Low-rank-SVGP-NP-update}{https://github.com/areferezaee/Low-rank-SVGP-NP-update}
Original source
This story was published by arXiv cs.CV and written by Arefeh Rezaei. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


