
AS
Alexandra Schild, Gerard de Melo
· 1 min read
ResearcharXiv cs.CV
Can't Find Waldo: Evaluating VLMs' Sensitivity to Image Resolution and Detail Level
arXiv:2609.31706v1 Announce Type: new
Abstract: Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high resolutions is lacking. We introduce a controlled evaluation framework that disentangles resolution-related performance degradation from task difficulty through semantics-preserving transformations. We propose two simple metrics: Area Under the Scaling Curve (AUSC), which quantifies scaling robustness independent of baseline accuracy, and Prediction Variance Score (PVS), which measures resolution-induced prediction instability. Through comprehensive experiments across 5 model families and 5 benchmarks, we identify three primary failure modes: (1) information loss from downsampling at vision token limits, (2) tokenization artifacts from patch boundary shifts and positional encoding fragility under non-standard aspect ratios, and (3) attention dilution as token counts increase. Our analysis reveals that even state-of-the-art models suffer from performance drops when processing high-resolution images, with degradation patterns varying systematically by architectural family. We provide actionable insights for model architecture design and data augmentation strategies to mitigate these limitations.
Original source
This story was published by arXiv cs.CV and written by Alexandra Schild, Gerard de Melo. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


