Hybrid Spatio-Temporal Feature Representation for Discriminative Video Summarization
DOI:
https://doi.org/10.70917/ijcisim-2026-4474Keywords:
Video Summarization, Feature Representation, Spatio-Temporal Modelling, Feature Fusion, PCA, Multi-Level FeaturesAbstract
The rapid growth of video data across domains such as surveillance, multimedia streaming, and social media platforms has necessitated the development of efficient video summarization techniques. A key factor influencing the performance of such systems is the quality of feature representation. Existing approaches often rely on limited or loosely integrated feature sets, resulting in redundancy and reduced discriminative capability.
This paper proposes a hybrid spatio-temporal feature representation framework that integrates contrast-based descriptors, object-level semantic features, and scene-level contextual representations into a unified feature space. The primary objective is to enhance feature discriminability and compactness by capturing complementary information across multiple semantic levels. To evaluate temporal consistency within the learned feature space, a standard Transformer encoder is employed as a generic sequence modelling component without architectural modification.
Experimental evaluation on the TVSum and SumMe datasets demonstrates that the proposed representation improves feature separability and contributes to enhanced summarization performance compared to conventional approaches. The results highlight the effectiveness of feature-centric design in improving discriminative capability while maintaining temporal consistency.