The paper “STAViS: Spatio-Temporal AudioVisual Saliency Network” was published at CVPR 2020, one of the leading venues in computer vision.
STAViS proposes a spatio-temporal audiovisual saliency model, combining visual and audio cues for video attention modelling. It is a strong multimedia research highlight for the group, connecting computer vision, audio-visual perception and deep learning.
- Authors: Antigoni Tsiami, Petros Koutras, Petros Maragos
- Venue: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020)
- Conference dates: 14-19 June 2020
- DOI: 10.1109/CVPR42600.2020.00482
- DBLP: bibliographic record
You may also like
-
TMLR 2026: prescriptive SVD-inspired attention via spectral energy retention
-
Towards Heritage World Models: from digital twins to predictive heritage systems
-
ICASSP 2026: representation-diverse self-supervision for bioacoustic learning
-
AISTATS 2026: process-tensor tomography of SGD and training memory
-
ICPRAM 2026: two-stage angular alignment for positive-unlabeled learning