Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
In International Conference on Pattern Recognition (ICPR) 2026 · Published
Abstract
The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such capabilities requires to approach several complex challenges. This work addresses the problem of human-object interaction anticipation in Egocentric Vision using Vision Large Language Models (VLLMs). We tackle key limitations in existing approaches by improving visual grounding capabilities through Set-of-Mark prompting and understanding user intent via the trajectory formed by the user's most recent gaze fixations. To effectively capture the temporal dynamics immediately preceding the interaction, we further introduce a novel inverse exponential sampling strategy for input video frames.
Experiments conducted on the egocentric dataset HD-EPIC demonstrate that our method surpasses state-of-the-art approaches for the considered task, showing its model-agnostic nature.
Publication record
- Venue
- International Conference on Pattern Recognition (ICPR) 2026
- Proceedings
- Pattern Recognition — Lecture Notes in Computer Science, Springer Nature Switzerland, Cham
- Pages
- 139–154
- DOI
- 10.1007/978-3-032-31404-8_10
- Published
- 3 August 2026 (online) · ISBN 978-3-032-31404-8
Citation
Materia, D., Ragusa, F., Farinella, G. M. (2026). Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos. In: De Marsico, M., et al. (eds) Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science. Springer, Cham, pp. 139–154.
@inproceedings{materia2026leveraging,
author = {Materia, Daniele and Ragusa, Francesco and Farinella, Giovanni Maria},
editor = {De Marsico, Maria and Ho, Tin Kam and Jurie, Frederic and Liu, Cheng-Lin and Lopresti, Daniel and Nystr{\"o}m, Ingela and Ogier, Jean-Marc and Ross, Arun and Wang, Liang},
title = {Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos},
booktitle = {Pattern Recognition},
series = {Lecture Notes in Computer Science},
publisher = {Springer Nature Switzerland},
address = {Cham},
year = {2026},
pages = {139--154},
isbn = {978-3-032-31404-8},
doi = {10.1007/978-3-032-31404-8_10}
}