Alzheimer's Disease Detection Using Linguistic Pattern Analysis of Low-Resource Kannada Transcripts
DOI:
https://doi.org/10.70917/ijcisim-2026-4390Keywords:
Alzheimer's Disease Detection, Natural Language Processing, Low-Resource Kannada NLP, DementiaBank Pitt Corpus, SVM-RBF, Cognitive Decline Screening, Linguistic Biomarkers, Speech TranscriptsAbstract
Alzheimer's disease (AD) is a progressive neurodegenerative disorder that impairs cognition, memory, and language ability, with global prevalence expected to surpass 150 million cases by 2050. Early detection remains critical for timely intervention and care planning. This paper proposes a novel NLP-based Alzheimer's detection framework that applies Natural Language Processing (NLP) and Machine Learning (ML) applied to speech transcripts in the low-resource Dravidian language Kannada. 484 DementiaBank Pitt Corpus transcripts are translated into Kannada, preprocessed with a language-aware CHAT-normalization pipeline, and a comprehensive 92-feature set spanning disfluency, lexical, syntactic, POS, semantic coherence, and information-content dimensions is extracted. Five ML classifiers are evaluated under 10-fold stratified cross-validation; SVM with an RBF kernel achieves the best performance (Accuracy: 81.0%, AUC-ROC: 87.5%, F1-Score: 80.5%), demonstrating the effectiveness of linguistic biomarkers for robust, non-invasive, and scalable Alzheimer's screening in language-impoverished settings.