Precision-Weighted Fusion of Fine-Tuned Text and Speech Encoders for Calibrated and Trustworthy Multimodal Emotion Recognition

Authors

  • Rajashree M Byalal Department of Computer Science & Engineering, Jain (Deemed –to-be-university), Bengaluru, India.
  • Sunanda Das Department of Computer Science & Engineering, Jain (Deemed –to-be-university), Bengaluru, India.
  • M Kumaresan Department of Computer Science & Engineering, Jain (Deemed –to-be-university), Bengaluru, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-3777

Keywords:

Multimodal emotion recognition, uncertainty quantification, model calibration, precision-weighted fusion, IEMOCAP, LoRA, selective prediction, soft labels

Abstract

Multimodal speech emotion recognition (SER) systems are typically optimized and reported solely in terms of accuracy, yet deployment in affect sensitive applications additionally requires calibrated confidence, the ability to abstain on unreliable inputs and uncertainty estimates that reflect the genuine ambiguity of emotional expression, how decision level fusion architectures differ along these trustworthiness dimensions on the IEMOCAP benchmark remains largely unquantified. We propose PWF-FT, a Precision Weighted Fusion framework in which a RoBERTa text encoder and a HuBERT speech encoder, adapted with low rank adaptation (LoRA) and learned attention pooling, each emit class logits together with an utterance level aleatoric log variance, and the two modality decisions are fused by inverse variance (precision) weights computed per utterance, training combines a heteroscedastic unimodal objective with a soft label distillation term derived from crowd annotation distributions  and the framework is benchmarked against text only, audio only, feature concatenation, fixed average and attention gated fusion under leave one session out (LOSO) evaluation with temperature scaling, expected calibration error (ECE), area under the risk coverage curve (AURC)  and selective prediction analyses. Fusion lifts weighted accuracy from 66.45% (text) and 58.02% (audio) to 71.13-73.14% across the four fusion variants, with attention gated fusion best at 73.14±1.79% WA (72.74±1.88% UA, 72.97±1.77% wF1) and PWF at 71.13±2.22% WA, after temperature scaling all fusion systems calibrate to ECE 0.037-0.052 and PWF is the only variant whose predictive entropy correlates positively and consistently with human soft label entropy (Spearman ρ = +0.286±0.029 across all five folds). We conclude that attention, concatenation and fixed fusion are statistically indistinguishable in accuracy once encoders are fine tuned (paired fold level tests, p ≥ 0.15), that precision weighting costs a small, quantified two WA points (p = 0.003) and that its practical value lies in interpretable, human aligned uncertainty that supports selective prediction and reliable deferral in human computer interaction, affective monitoring, and clinical screening pipelines.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-27

How to Cite

Rajashree M Byalal, Sunanda Das, & M Kumaresan. (2026). Precision-Weighted Fusion of Fine-Tuned Text and Speech Encoders for Calibrated and Trustworthy Multimodal Emotion Recognition. International Journal of Computer Information Systems and Industrial Management Applications, 18(11s), 629–648. https://doi.org/10.70917/ijcisim-2026-3777

Issue

Section

Original Articles