TE-CUME: Uncertainty-Aware Tri-Modal Video Retrieval Representation Learning
DOI:
https://doi.org/10.70917/ijcisim-2026-5179Keywords:
Video Representation Learning, Uncertainty Quantification, Multimodal Fusion, Dense Video Captioning, Monte Carlo Dropout, Calibration, Video-Text RetrievalAbstract
Automatic Content Recognition (ACR) systems deployed in smart TVs rely on multimodal encoders that produce deterministic fingerprint embeddings, offering no mechanism to express how trustworthy a particular identification is under varying broadcast degradation. This paper introduces TE-CUME (Tri-Encoder Cross-Modal Uncertainty Fusion Model), a probabilistic tri-modal encoder that equips each modality stream — visual (ViViT-S), audio (BEATs), and on-screen text (TrOCR) — with Monte Carlo Dropout and heteroscedastic output heads, producing per-modality Gaussian posteriors over the content embedding space. A minimum-variance Product-of-Experts fusion stage combines these posteriors through precision weighting, provably minimising the fused estimate’s variance. We propose UAT-InfoNCE, a precision-weighted contrastive objective that adaptively modulates gradient contributions by the model’s own confidence, together with pairwise cross-modal disagreement terms (,,) that serve as zero-shot diagnostics for broadcast failure modes such as A/V desynchronisation, dubbed content, and advertisement splices. Evaluation on ActivityNet Captions focuses primarily on video-text retrieval as a direct encoder-quality benchmark across the full 4,917-video validation set (R@, mAP, ECE), complemented by a conceptual architectural specification showing how TE-CUME serves as a modular Video Feature Extractor (VFE) replacement for dense video captioning pipelines (PDVC, Vid2Seq) while propagating calibrated per-event confidence scores. Experimental results demonstrate that predicted epistemic uncertainty correlates positively with degradation severity (Spearman,), and that TE-CUME achieves competitive retrieval performance (R@1 = 43.1 on ActivityNet) while uniquely providing per-query calibrated confidence that enables downstream systems to act conservatively on unreliable representations.