
AC
Anu Chowdhury, Bin Wu, Hossein A. Rahmani, Emine Yilmaz
· 1 min read
ResearcharXiv cs.CL
Automatic multimodal UX improvement recommendations from LLM agent user simulations
arXiv:2609.22971v1 Announce Type: new
Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.
Original source
This story was published by arXiv cs.CL and written by Anu Chowdhury, Bin Wu, Hossein A. Rahmani, Emine Yilmaz. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


