Exploring Global Diversity and Local Context for Video Summarization

Pan, Yingchao; Huang, Ouhan; Ye, Qinghao; Li, Zhongjin; Wang, Wenjiang; Li, Guodun; Chen, Yuxing

Computer Science > Computer Vision and Pattern Recognition

arXiv:2201.11345 (cs)

[Submitted on 27 Jan 2022 (v1), last revised 27 Mar 2022 (this version, v2)]

Title:Exploring Global Diversity and Local Context for Video Summarization

Authors:Yingchao Pan, Ouhan Huang, Qinghao Ye, Zhongjin Li, Wenjiang Wang, Guodun Li, Yuxing Chen

View PDF

Abstract:Video summarization aims to automatically generate a diverse and concise summary which is useful in large-scale video processing. Most of the methods tend to adopt self-attention mechanism across video frames, which fails to model the diversity of video frames. To alleviate this problem, we revisit the pairwise similarity measurement in self-attention mechanism and find that the existing inner-product affinity leads to discriminative features rather than diversified features. In light of this phenomenon, we propose global diverse attention which uses the squared Euclidean distance instead to compute the affinities. Moreover, we model the local contextual information by novel local contextual attention to remove the redundancy in the video. By combining these two attention mechanisms, a video SUMmarization model with Diversified Contextual Attention scheme is developed, namely SUM-DCA. Extensive experiments are conducted on benchmark data sets to verify the effectiveness and the superiority of SUM-DCA in terms of F-score and rank-based evaluation without any bells and whistles.

Comments:	Accepted by IEEE Access
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2201.11345 [cs.CV]
	(or arXiv:2201.11345v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2201.11345

Submission history

From: Qinghao Ye [view email]
[v1] Thu, 27 Jan 2022 06:56:01 UTC (3,010 KB)
[v2] Sun, 27 Mar 2022 15:50:58 UTC (6,023 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring Global Diversity and Local Context for Video Summarization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring Global Diversity and Local Context for Video Summarization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators