Multimodal Representation Learning and Visual Analytics for Modeling the Formal Visual Language of Contemporary Inner Mongolian Painting

Authors

  • Shuzhen Yan School of Fine Arts and Design, Jining Normal University, Ulanqab, Inner Mongolia, 830000, China

DOI:

https://doi.org/10.70917/ijcisim-2026-2821

Keywords:

Digital Humanities; Art Style Analysis; Multimodal Embedding; Visual Analytics

Abstract

This paper presents a computer vision-driven multimodal representation learning framework for modeling the formal visual language of contemporary Inner Mongolian painting. A corpus of approximately 1,200 digitized paintings is curated together with structured metadata collected from museums, regional galleries, exhibition catalogues, and artist archives. To move beyond descriptive digital archiving, we formulate artwork analysis as a joint image–metadata learning problem. Specifically, painting images are encoded by a deep visual backbone, while artwork-level metadata, including artist, year, genre, subject matter, medium, and descriptive tags, are encoded into semantic representations. The two modalities are projected into a shared visual-semantic latent space through contrastive alignment and further integrated by a metadata-aware fusion module. The resulting joint representations are evaluated on style classification, semantic retrieval, unsupervised clustering, and visual analytics. Compared with an image-only CNN baseline, the proposed multimodal framework improves style classification accuracy from 78% to 88%, semantic consistency from 65% to 80%, feature extraction accuracy from 70% to 85%, and Precision@10 from 70% to 82%. Qualitative analyses based on confusion matrices, embedding visualization, radar plots, and theme co-occurrence graphs further show that the learned representation better separates stylistic categories while preserving culturally meaningful motifs and thematic relations. The study demonstrates that multimodal representation learning can provide an interpretable computational pipeline for regional artwork analysis and offers a data-driven bridge between computer vision and digital humanities.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-08

How to Cite

Shuzhen Yan. (2026). Multimodal Representation Learning and Visual Analytics for Modeling the Formal Visual Language of Contemporary Inner Mongolian Painting. International Journal of Computer Information Systems and Industrial Management Applications, 18, 11. https://doi.org/10.70917/ijcisim-2026-2821

Issue

Section

Original Articles