ICPR 2026 conference logo

Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos

In International Conference on Pattern Recognition (ICPR) 2026 · Published

Abstract

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such capabilities requires to approach several complex challenges. This work addresses the problem of human-object interaction anticipation in Egocentric Vision using Vision Large Language Models (VLLMs). We tackle key limitations in existing approaches by improving visual grounding capabilities through Set-of-Mark prompting and understanding user intent via the trajectory formed by the user's most recent gaze fixations. To effectively capture the temporal dynamics immediately preceding the interaction, we further introduce a novel inverse exponential sampling strategy for input video frames.

Experiments conducted on the egocentric dataset HD-EPIC demonstrate that our method surpasses state-of-the-art approaches for the considered task, showing its model-agnostic nature.

Publication record

Venue
International Conference on Pattern Recognition (ICPR) 2026
Proceedings
Pattern Recognition — Lecture Notes in Computer Science, Springer Nature Switzerland, Cham
Pages
139–154
DOI
10.1007/978-3-032-31404-8_10
Published
3 August 2026 (online) · ISBN 978-3-032-31404-8

Citation

Materia, D., Ragusa, F., Farinella, G. M. (2026). Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos. In: De Marsico, M., et al. (eds) Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science. Springer, Cham, pp. 139–154.

@inproceedings{materia2026leveraging,
  author    = {Materia, Daniele and Ragusa, Francesco and Farinella, Giovanni Maria},
  editor    = {De Marsico, Maria and Ho, Tin Kam and Jurie, Frederic and Liu, Cheng-Lin and Lopresti, Daniel and Nystr{\"o}m, Ingela and Ogier, Jean-Marc and Ross, Arun and Wang, Liang},
  title     = {Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos},
  booktitle = {Pattern Recognition},
  series    = {Lecture Notes in Computer Science},
  publisher = {Springer Nature Switzerland},
  address   = {Cham},
  year      = {2026},
  pages     = {139--154},
  isbn      = {978-3-032-31404-8},
  doi       = {10.1007/978-3-032-31404-8_10}
}