Mining Mythic and Ideological Discourse in Science Fiction Literature: A Computational NLP Approach
DOI:
https://doi.org/10.70917/ijcisim-2026-4283Keywords:
Natural Language Processing, Ideological Discourse Detection, Named Entity Recognition, Topic Modelling, Sentiment Analysis, Mythic Discourse, Science Fiction Corpus, BERT, LDA, Dune TrilogyAbstract
In this paper, we proposed a multi-layer Natural Language Processing (NLP) system which is able to detect, classify and track ideologically and mythically the patterns of discourse in large bodies of literary texts in time. By using Named Entity Recognition (NER), custom Myth-Entity Tagging, Latent Dirichlet Allocation (LDA) topic modelling, and BERT-based sentiment analysis, the proposed end-to-end pipeline is able to quantify the propagation and emotional flow of prophetic, messianic and simulacral language within long fictional texts. In the present study, the Dune trilogy written by the author Frank Herbert (Dune (1965), Dune Messiah (1969), and Children of Dune (1976)) is used as one of the main corpora, with approximately 380,000 tokens, as a structural tool of the imperial ideology in the context of engineered myth-making. The experimental results confirm the changes that the TF-IDF of the terms related to the markers of the myth undergo, given that they vary from one novel to another, and that the NER has identified seven categories of mythic entities, with the total number of instances being 1461, while the sentiment scores provided by BERT consistently decrease sentiment in prophetic discourse throughout the course of the novels, all of which lead to the hypothesis that the act of liberating a myth becomes simulacral control along its course. This framework can be generalized across any literary body of work that is ideologically dense, and offers a repeatable computational approach for digital humanities research in the field of NLP, ideology critique and science fiction.