Eliminating Single Point Failures through Probabilistic Consensus, ZooKeeper Coordination, and Kafka-Style Replication in HDFS Storage Systems

Authors

  • Rohini D. Dhage VMV Commerce, JMT Arts & JJP Science College, Nagpur, Maharashtra, India.
  • Vaibhav R. Bhedi VMV Commerce, JMT Arts & JJP Science College, Nagpur, Maharashtra, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-4830

Keywords:

Single Point Failure, HDFS High Availability, Probabilistic Consensus, ZooKeeper Coordination, Kafka Replication, NameNode Failover, Metadata Resilience, Fault Tolerance

Abstract

Although a significant weakness in HDFS-like storage architectures is that the system still relies on single points of failure due to the fact that most metadata leadership is currently controlled by a single active NameNode with passive NameNodes waiting for some form of signal upon an active node's failure; we are presenting here, a Probabilistic Consensus-Driven Active-Standby Controller (PC-ASC) as a model to remove these weak linkages using Bayesian health beliefs, majoritarian quorums to validate, ZooKeeper-based coordination, and Kafka-style data replication of metadata logs. The PC-ASC also includes an evaluation framework that continually monitors and measures the performance of heartbeat quality, metadata latency, the state of the network topology, and the values of the vector representing failure probabilities, to generate trust aware role assignment decisions and decision-based failover signals. We have utilized a controlled emulation data set consisting of metadata requests, heartbeat trace information, injected crash events, and duplicated edit-log record sets to perform comparative evaluations of the proposed approach relative to the single-name-node architecture, the traditional HA-ZKFC/QJM architecture, and the raft-like metadata clustering baseline approaches. Our results indicate fail-over times were reduced from 15.2 seconds down to 6.0 seconds, our available time was increased from 99.59% to 99.999%, and our metadata throughput improved by as much as 3.04 times compared to when we had only a single NameNode controlling all operations. Ablations studies confirm that consensus mechanisms were important to achieving improvements, trust scoring was important to achieving improvements, and the use of log-based recovery was important to achieve improvements in process.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-17

How to Cite

Rohini D. Dhage, & Vaibhav R. Bhedi. (2026). Eliminating Single Point Failures through Probabilistic Consensus, ZooKeeper Coordination, and Kafka-Style Replication in HDFS Storage Systems. International Journal of Computer Information Systems and Industrial Management Applications, 18(17s), 1336–1350. https://doi.org/10.70917/ijcisim-2026-4830

Issue

Section

Original Articles