A Lightweight, Transparent Machine Learning System for Real-Time Detection of Fraudulent Online Job Postings

Authors

  • Syed Umraz Department of Computer Science and Engineering, PDA College of Engineering, Kalaburagi, Karnataka, India
  • Radha B.K Department of Computer Science and Engineering, PDA College of Engineering, Kalaburagi, Karnataka, India
  • Anuradha T Department of Computer Science and Engineering, PDA College of Engineering, Kalaburagi, Karnataka, India
  • Megha.M. Shete Department of Computer Science and Engineering, PDA College of Engineering, Kalaburagi, Karnataka, India

DOI:

https://doi.org/10.70917/ijcisim-2026-5736

Keywords:

online recruitment fraud, fake job posting detection, TF-IDF vectorization, K-Nearest Neighbors, Random Forest, , job-scam classification, user-in-the-loop retraining

Abstract

Background. Online recruitment fraud — fake job postings that solicit upfront fees, harvest personal data, or lure applicants into unpaid work — is a persistent problem on job boards and professional networks. Prior work has shown that classical text-classification pipelines (TF-IDF plus tree-based or linear classifiers) can detect such postings with high accuracy on large curated corpora, but most published systems are evaluated offline and are not integrated into a deployable, user-facing product with authentication, audit history, and a data-collection loop. Methods. We present a Flask-based web system that classifies job postings as legitimate or fraudulent using TF-IDF vectorization (200 features, unigrams and bigrams, English stop-word removal) of the concatenated job title, company name, description, and salary text, followed by one of three switchable scikit-learn classifiers — K-Nearest Neighbors (k=5, distance-weighted, default), Random Forest (100 estimators, max depth 10), or Logistic Regression (max_iter 1000) — trained on a stratified 80/20 split (random_state=42). The system persists every prediction and every retraining run to a SQLite database, exposes a Selenium-based LinkedIn extraction feature so operators can grow the training set from live postings, and lets a logged-in user retrain the model on demand from an uploaded CSV. Results. On the project's 34-row labeled dataset (24 legitimate, 10 fraudulent; 27 training / 7 test rows, stratified, random_state=42), the default KNN configuration achieved 85.71% accuracy, 66.67% precision, 100% recall, and an F1-score of 80.00%. Random Forest achieved 100% on all four metrics on the same 7-row test fold; Logistic Regression achieved 85.71% accuracy, 100% precision, 50% recall, and an F1-score of 66.67%. Given the small sample size and the feature-count-to-training-row ratio (200 TF-IDF features from 27 documents), these figures should be read as a reproducible snapshot of current behavior, not a generalization guarantee. Conclusion. A minimal, fully transparent TF-IDF plus classical-classifier pipeline is sufficient to build a working, auditable fake-job-posting detector inside a real deployable product. The system's main contribution is the end-to-end integration — authentication, prediction logging, on-demand retraining, and a live data-collection path — rather than a novel classification algorithm. Before any accuracy claim is generalized, the dataset must be substantially expanded (ideally validated against a public benchmark such as EMSCAD or the derived Kaggle "Fake Job Postings" corpus) and evaluated with cross-validation rather than a single held-out split.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-04

How to Cite

Syed Umraz, Radha B.K, Anuradha T, & Megha.M. Shete. (2026). A Lightweight, Transparent Machine Learning System for Real-Time Detection of Fraudulent Online Job Postings. International Journal of Computer Information Systems and Industrial Management Applications, 18(18s), 1528–1543. https://doi.org/10.70917/ijcisim-2026-5736

Issue

Section

Original Articles