SyncAI.news, a Varaisys broadcasting
GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction
SZ

Sheng Zhao, Weikai Lin, Yuhao Zhu

· 1 min read

ResearcharXiv cs.CV

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

arXiv:2609.38519v1 Announce Type: new Abstract: Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.

Original source

This story was published by arXiv cs.CV and written by Sheng Zhao, Weikai Lin, Yuhao Zhu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News