Socio-Phonosemantic Digitization of Blaan Cultural Terminologies Through an Empirical Dataset and Ethical Data Governance Framework
DOI:
https://doi.org/10.70917/ijcisim-2026-3373Keywords:
Artificial Intelligence, Indigenous Data Sovereignty, Natural Language Processing, Ethical Data Governance, Low-Resource Languages, Blaan languageAbstract
This study aims to document the socio-phonosemantic terminologies of the Blaan language to structure an empirical dataset foundational for culturally responsive artificial intelligence (AI) and digital information systems. Employing a descriptive-qualitative research design through ethnographic methodologies, the study executed rigorous data acquisition and validation. Informants were selected via purposive and convenience sampling from within the Blaan community to ensure cultural authenticity and structural integrity. Findings present a structured dataset revealing distinct linguistic features, including the presence of the schwa sound, frequent use of the phonemes /f/ and /v/, and multiple consonant clusters (e.g., /bl/, /fn/, /kw/). These phonetic nuances present specific computational challenges that require bespoke natural language processing (NLP) algorithmic training rather than rigid, pre-existing templates. Furthermore, the documentation of deep cultural significance, such as sensitive terminologies like ki and traditional de jure laws governed by tribal leaders, demonstrates the critical necessity of embedding Role-Based Access Control (RBAC) and decentralized administrative hierarchies into digital repositories. Ultimately, this research bridges localized indigenous knowledge with modern digital tools, establishing an ethical data governance framework. By aligning traditional customs with Indigenous Data Sovereignty, the study provides a vital blueprint for ensuring that technological integration protects, rather than exploits, marginalized ethnolinguistic identities.