Federated Learning-Based Privacy-Preserving Malware Detection Framework Using Multi- Source Cybersecurity Datasets

Bachu Anusha

The increasing prevalence and sophistication of malware attacks in distributed and heterogeneous computing environments necessitate advanced detection mechanisms that ensure both high accuracy and data privacy. Traditional centralized malware detection systems are limited by privacy concerns, regulatory constraints, and the inability to effectively utilize data from multiple sources. To address these challenges, this study proposes a Privacy-Preserving Federated Learning (PPFL) framework for malware detection that enables collaborative model training across distributed clients without sharing raw data. The proposed approach integrates deep neural network-based local models with Federated Averaging for global aggregation, while incorporating Differential Privacy and Secure Aggregation mechanisms to protect sensitive information during communication. The framework is evaluated using multi-source cybersecurity datasets, including CCCS-CIC-AndMal-2020 and IoT-23, capturing both static and dynamic malware features. Experimental results demonstrate that the model achieves a detection accuracy of 96.42% after 50 communication rounds, with balanced precision, recall, and F1- score, indicating robust and reliable performance. The findings highlight the effectiveness of federated learning in enabling secure, scalable, and high-performance malware detection, making it a viable solution for real-world cybersecurity applications.
PDF