Multimodal Deep Representation Learning for Nutritional Status Prediction in Children Under Five
DOI:
https://doi.org/10.70917/ijcisim-2026-5334Keywords:
nutritional risk prediction, multimodal deep learning, variational autoencoder, transformer encoder, child health assessment, under-five nutritionAbstract
Nutritional risk among children under five years should be identified early to minimize the developmental, health and cognitive impairments in the long run. The assessment of the nutritional status is however usually confronted by disaggregated and diverse sources of data such as anthropometric data, dietary behavior records, social economic data and clinical health history. The proposed study is a multimodal deep representation learning model, which will learn structured and unstructured child health data to predict the nutrition status correctly. The model encodes structured clinical and anthropometric features into latent representation with a Variational Autoencoder and encodes time and contextual patterns of diet and feeding behavior with a Transformer encoder. The cross-modal fusion mechanism takes these latent embeddings and unites them into a single representation of child health with which nutritional risk can be classified. The suggested approach makes it possible to do solid data-driven nutritional profiling and facilitate detecting children, who are at moderate or severe nutritional risk, in their early-stage. Experimental analysis indicates a better classification accuracy and stability compared to base machine learning and single-stream deep models. This framework provides a computational framework scaling method on enhancing the pediatric nutritional surveillance and targeted intervention planning.