Fedagentshield: An Adaptive Federated Reinforcement Learning Framework for Privacy-Preserving Autonomous Cyber Defence
DOI:
https://doi.org/10.70917/ijcisim-2026-4511Keywords:
Federated Reinforcement Learning, Federated Learning, Autonomous Cyber Defence, Proximal Policy Optimisation, Privacy-Preserving Learning, Adaptive Federated Policy Aggregation, Network Intrusion DetectionAbstract
The growing sophistication and frequency of cyber threats, including Advanced Persistent Threats, ransomware, zero-day exploits, and AI-assisted attacks, highlight the limitations of traditional centralised cybersecurity systems in scalability, adaptability, and data privacy. Federated Learning (FL) and Reinforcement Learning (RL) have shown significant promise for distributed learning and intelligent decision-making across a variety of applications. However, their synergistic combination for privacy-preserving autonomous cyber defence is still an open research issue. The paper proposes FedAgentShield, an adaptive Federated Reinforcement Learning (FRL) framework that enables multiple distributed organisations to collaboratively learn effective cyber defence policies without sharing cybersecurity data. The proposed scheme features a PPO-based autonomous defence agent and a federated learning architecture for decentralised policy learning in distributed cyber environments. The framework can enhance the quality of global policy updates by introducing an Adaptive Federated Policy Aggregation (AFPA) mechanism. The AFPA mechanism forms aggregation weights by leveraging multiple client performance indicators, including detection performance, reinforcement learning reward, response efficiency, and privacy confidence, rather than conventional data-size-based aggregation. We evaluate the proposed framework in a custom cyber defence environment with the NF-UNSW-NB15 benchmark dataset. Findings reveal that FedAgentShield enables collaborative policy learning while protecting user privacy and enhancing overall policy quality through adaptive weighting. The proposed structure is a scalable, adaptive, and privacy-preserving approach for next-generation autonomous cyber defence in distributed network environments.