Deep heterogeneity learning for cross-city transit forecasting: a differentially private federated framework with mixture-of-experts and seasonal decomposition
Introduction Accurate prediction of transit flows is fundamental to optimizing intelligent transportation systems; however, centralized forecasting is frequently obstructed by heterogeneous, Non-Independent and Identically Distributed (Non-IID) cross-city data and stringent data privacy regulations. Methods We propose X-FedFormer, a novel framework integrating Federated Learning (FL) with Differential Privacy (DP) and a deep learning architecture combining a Mixture-of-Experts (MoE) mechanism with a Seasonal-Trend Decomposition module. The framework is evaluated on a statistically validated synthetic dataset faithfully simulating realistic inflow and outflow patterns across ten diverse urban environments (90 days of hourly records, 30 routes per city, 64,800 observations per city). Results X-FedFormer significantly outperforms state-of-the-art federated baselines including FedProx, achieving an aggregate coefficient of determination of 0.922 and a mean absolute error (MAE) of 7.93 passengers across all participating cities. A Wilcoxon signed-rank test confirms statistical significance over the strongest baseline (p = 0.018). Ablation studies confirm that the MoE and seasonal decomposition modules reduce forecasting error by approximately 11% and 16%, respectively, compared to standard architectures. Discussion The model maintains high predictive utility even under strict differential privacy guarantees (ε ≈ 2), establishing a viable privacy-utility operating point for practical deployment. These findings present a scalable, robust solution for urban computing that effectively balances algorithmic performance with data sovereignty in smart city applications.
- Research Article
10
- 10.1109/access.2024.3523909
- Jan 1, 2025
- IEEE Access
The problem of data privacy protection in the information age deserves people’s attention. As a distributed machine learning technology, federated learning can effectively solve the problem of privacy security and data silos. Differential privacy(DP) technology is applied in federated learning(FL). By adding noise to raw data and model parameters, it can further enhance the degree of data privacy protection. Over the years, differential privacy technology based on federated learning framework has been developed, which is divided into central differential privacy federated learning(CDPFL) and local differential privacy federated learning(LDPFL). Although differential privacy may reduce the accuracy and convergence of federated learning models while protecting data privacy, researchers have proposed a variety of optimization methods to balance privacy protection and model performance. This paper comprehensively expounds the research status of differential privacy techniques based on the federated learning framework, first providing detailed introductions to federated learning and differential privacy technologies, and then summarizing the development status of two types of federated learning differential privacy(DPFL) techniques respectively; for CDPFL, the paper divides the discussion into first proposal of CDP and typical application examples, the impact of Gaussian mechanisms on model accuracy, optimization based on asynchronous differential privacy, and insights from other scholars; for LDPFL, the paper divides the discussion into first proposal of LDP and typical application examples, processing multidimensional data and improving model accuracy, existing methods and optimization for reducing communication costs, balancing privacy protection and data usability, LDPFL based on the Shuffle model, and insights from other scholars; following this, the paper addresses and summarizes the unique challenges introduced by incorporating differential privacy into federated learning and proposes solutions; finally, based on a summary of existing optimization techniques, the paper outlines future directions and specifically discusses three research ideas for enhancing the optimization effects of federated differential privacy: advanced optimization strategies combining Bayesian methods and the Alternating Direction Method of Multipliers (ADMM), integrating lattice homomorphic encryption techniques from cryptography to achieve more efficient differential privacy protection in federated learning, and exploring the application of zero-knowledge proof techniques in federated learning for privacy protection.
- Research Article
2
- 10.1109/access.2025.3647561
- Jan 1, 2025
- IEEE Access
Medical imaging plays a crucial role in the diagnosis of various diseases, including cancer, cardiovascular conditions, and respiratory disorders such as asthma, pneumonia, and even COVID-19. Nowadays, Deep Learning (DL) has demonstrated remarkable accuracy in disease detection and diagnosis using medical imaging such as X-rays, CT scans, and MRIs. However, DL models mostly require access to sensitive patient data not only to capture the images but also to store them in the data cloud, which raises major privacy concerns. To solve this problem, several techniques such as Differential Privacy (DP) and Federated Learning (FL), have been developed. These methods allow models to learn useful patterns without exposing personal details. Building on these developments, the key gap this review targets is the lack of a recent, side-by-side analysis that connects medical tasks, DL architectures, FL designs, and DP budgets with real-world performance and resilience against privacy attacks. Motivated by this gap, we conducted a systematic review of the literature (SLR) of studies published between 2023 and 2025 that used DL principles with privacy-preserving methods for medical diagnosis. Utilizing the PRISMA framework, we reviewed 23 peer-reviewed articles and grouped them into four categories: included COVID-19, skin lesions, Alzheimer’s disease, and other conditions such as diabetic retinopathy, kidney disease, and heart problems. We sort results by modality, model family, FL topology, DP mechanism, and reported threat model. Furthermore, the review significantly represents that CNN-based models with DP and FL methods are the most common choices, and developed pipelines with these models often achieve accuracy above 90%. At the same time, the review found gaps such as limited real-world testing, a lack of diverse datasets, and underuse of advanced DP approaches. Based on these results, we recommend building larger and more varied datasets, improving privacy-utility trade-offs, and testing models in clinical environments.
- Research Article
94
- 10.3233/his-220006
- May 31, 2022
- International Journal of Hybrid Intelligent Systems
Federated learning (FL) refers to a system in which a central aggregator coordinates the efforts of several clients to solve the issues of machine learning. This setting allows the training data to be dispersed in order to protect the privacy of each device. This paper provides an overview of federated learning systems, with a focus on healthcare. FL is reviewed in terms of its frameworks, architectures and applications. It is shown here that FL solves the preceding issues with a shared global deep learning (DL) model via a central aggregator server. Inspired by the rapid growth of FL research, this paper examines recent developments and provides a comprehensive list of unresolved issues. Several privacy methods including secure multiparty computation, homomorphic encryption, differential privacy and stochastic gradient descent are described in the context of FL. Moreover, a review is provided for different classes of FL such as horizontal and vertical FL and federated transfer learning. FL has applications in wireless communication, service recommendation, intelligent medical diagnosis system and healthcare, which we review in this paper. We also present a comprehensive review of existing FL challenges for example privacy protection, communication cost, systems heterogeneity, unreliable model upload, followed by future research directions.
- Research Article
128
- 10.3390/healthcare12242587
- Dec 22, 2024
- Healthcare (Basel, Switzerland)
Federated learning (FL) is revolutionizing healthcare by enabling collaborative machine learning across institutions while preserving patient privacy and meeting regulatory standards. This review delves into FL's applications within smart health systems, particularly its integration with IoT devices, wearables, and remote monitoring, which empower real-time, decentralized data processing for predictive analytics and personalized care. It addresses key challenges, including security risks like adversarial attacks, data poisoning, and model inversion. Additionally, it covers issues related to data heterogeneity, scalability, and system interoperability. Alongside these, the review highlights emerging privacy-preserving solutions, such as differential privacy and secure multiparty computation, as critical to overcoming FL's limitations. Successfully addressing these hurdles is essential for enhancing FL's efficiency, accuracy, and broader adoption in healthcare. Ultimately, FL offers transformative potential for secure, data-driven healthcare systems, promising improved patient outcomes, operational efficiency, and data sovereignty across the healthcare ecosystem.
- Research Article
7
- 10.1109/tdsc.2023.3341788
- Jul 1, 2024
- IEEE Transactions on Dependable and Secure Computing
Federated Learning (FL) is a collaborative learning framework that enables edge devices to collaboratively learn a global model while keeping raw data locally. Although FL avoids leaking direct information from local datasets, sensitive information can still be inferred from the shared models. To address the privacy issue in FL, differential privacy (DP) mechanisms are leveraged to provide formal privacy guarantee. However, when deploying FL at the wireless edge with over-the-air computation, ensuring client-level DP faces significant challenges. In this paper, we propose a novel wireless FL scheme called private federated edge learning with sparsification (PFELS) to provide client-level DP guarantee with intrinsic channel noise while reducing communication and energy overhead and improving model accuracy. The key idea of PFELS is for each device to first compress its model update and then adaptively design the transmit power of the compressed model update according to the wireless channel status without any artificial noise addition. We provide a privacy analysis for PFELS and prove the convergence of PFELS under general non-convex and non-IID settings. Experimental results show that compared with prior work, PFELS can improve the accuracy with the same DP guarantee and save communication and energy costs simultaneously.
- Research Article
- 10.3389/fdgth.2026.1691088
- Jan 1, 2026
- Frontiers in digital health
Federated learning (FL) has the potential to boost deep learning in neuroimaging but is rarely deployed in real-world scenarios, where its true potential lies. We propose FLightcase, a new FL toolbox tailored for brain research, and evaluate it on a real-world FL network to predict the cognitive status in patients with multiple sclerosis (MS) from brain magnetic resonance imaging (MRI). We first trained a DenseNet neural network to predict age from T1-weighted brain MRI on three open-source datasets: IXI (586 images), SALD (491 images), and CamCAN (653 images). These were distributed across the three centres in our FL network: Brussels (BE), Greifswald (DE), and Prague (CZ). We benchmarked this federated model with a centralised version. The best-performing brain age model was then fine-tuned to predict performance on the symbol digit modalities test (SDMT) of patients with MS (Brussels: 96 images, Greifswald: 756 images, Prague: 2,424 images). Shallow transfer learning (TL) was compared with deep transfer learning, in which weights were updated either in the last layer or across the entire network, respectively. Federated training outperformed centralised training, predicting age with a mean absolute error (MAE) of 6.08 versus 7.02. Federated training yielded Pearson correlations (all p < .001) between true and predicted age of0.88 (IXI, Brussels), 0.91 (SALD, Greifswald), and 0.93 (CamCAN, Prague). Fine-tuning of the centralised model to SDMT was most successful with a deep TL paradigm (MAE = 9.19) compared to shallow TL (MAE = 11.05). Across Brussels, Greifswald, and Prague, deep TL predicted SDMT with MAEs of 10.71, 9.67, and 8.98, respectively, and yielded Pearson correlations between true and predicted SDMT of.25 (p = 0.282), 0.40 (p < 0.001), and 0.50 (p < 0.001). Real-world federated learning using FLightcase is feasible for neuroimaging research in MS, enabling access to large MS imaging databases without sharing data. The federated SDMT-decoding model is promising and could be improved in the future by adopting FL algorithms that address the non-IID data issue and consider other imaging modalities. We hope our detailed real-world experiments and open-source distribution of FLightcase will prompt researchers to move beyond simulated FL environments.
- Research Article
- 10.1371/journal.pone.0342454
- Jan 1, 2026
- PloS one
In smart grids, data collection is carried out through smart meters and devices of the Internet of Things, which are installed in the home, allowing to predict the demand for electricity and optimize the distribution of energy. Although the smart grids improve efficiency of operations for end users, they simultaneously present pronounced challenges regarding user privacy and security at the system level. In the context of conventional centralized machine learning, paradigms risk breaching the raw data of consumers, while decentralized paradigms often lack strong mechanisms for verifying identity or ensuring traceability. Existing federated learning systems often lack client level differential privacy, secure aggregation, and decentralized identity protection, leaving them vulnerable to privacy leakage and inference attacks. Blockchain based solutions typically expose model updates or use single layer identifiers. This paper introduces a secure and privacy preserving architecture that combines a dual layer blockchain architecture, federated learning (FL) and central differential privacy (DP) to thoroughly solve these challenges. The proposed system includes a dual layer blockchain system that ensures secure and tamper resistant logging of client interactions and protects client identities by storing salted cryptographic hashes. This design provides both traceability and anonymity, and thus maintains the integrity of participation while obfuscating sensitive identifiers. Privacy is guaranteed by storing raw data in client devices and sending only model updates for central aggregation. At the server side, Gaussian noise is added to the aggregated model parameters to achieve central DP, so as to reduce the risks of inference attacks on user data. Implementation of the proposed framework was performed based on Flower to test the PRECON (Pakistan Residential Electricity CONsumption) dataset, which consists of real-world household electricity consumption data. Multiple machine learning models were benchmarked and out of all the models, Random Forest performed best with the performance metrics of Mean Absolute Error (MAE) of 0.153, Mean Absolute Percentage Error (MAPE) of 0.085 and Mean Squared Error (MSE) of 0.143. The results showed that the proposed framework improved data privacy, preserved the forecasting accuracy and security in smart grid environments.
- Research Article
260
- 10.1016/j.seta.2022.102987
- Dec 31, 2022
- Sustainable Energy Technologies and Assessments
Federated learning for smart cities: A comprehensive survey
- Research Article
40
- 10.1016/j.rineng.2024.102773
- Aug 24, 2024
- Results in Engineering
A comprehensive survey on load forecasting hybrid models: Navigating the Futuristic demand response patterns through experts and intelligent systems
- Research Article
47
- 10.1016/j.future.2023.10.013
- Oct 31, 2023
- Future Generation Computer Systems
FederatedTrust: A solution for trustworthy federated learning
- Research Article
4
- 10.1007/s44196-025-00829-0
- Nov 17, 2025
- International Journal of Computational Intelligence Systems
The rapid proliferation of smart cities has led to an unprecedented generation of sensitive data from interconnected infrastructures, such as healthcare, transportation, energy, and surveillance systems. Ensuring data privacy and security while enabling real-time data analytics remains a critical challenge as traditional centralized processing methods are vulnerable to cyber threats and regulatory constraints. While Federated Learning (FL) offers a decentralized alternative by keeping data local, it remains susceptible to privacy breaches during model updates. Homomorphic Encryption (HE) has emerged as a promising solution for securing computations on encrypted data, but its high computational overhead hinders its applicability in real-time scenarios. To address these limitations, this paper proposes Scalable Privacy-Preserving Federated Learning with Efficient Homomorphic Encryption (SPP-FLHE)—a novel framework that integrates FL with optimized HE techniques and Differential Privacy (DP). The proposed framework reduces computational and communication overhead while ensuring robust privacy protection. Our key contributions include: (1) an enhanced HE scheme that minimizes encryption-related latency, making FL more feasible for large-scale deployments; (2) a dynamic DP noise addition mechanism that strengthens privacy without significantly degrading model accuracy; and (3) the introduction of model compression and gradient sparsification techniques to reduce communication overhead, improving scalability for smart city applications. Extensive experiments on real-world smart city datasets demonstrate that SPP-FLHE achieves 92.6% accuracy, while reducing data transmission by 40% and latency by 43% compared to conventional FL-HE methods. These findings highlight the framework’s practicality for real-time smart city applications, enabling secure and efficient data analytics without compromising privacy. The proposed model represents a significant advancement in privacy-preserving machine learning, offering a scalable and computationally efficient approach for securing smart city infrastructures.
- Research Article
- 10.52710/cfs.363
- Feb 13, 2025
- Computer Fraud and Security
With the advancement of information technology, data security and user privacy protection have become paramount. To achieve efficient privacy protection in a federated learning environment, a differential privacy algorithm is designed using the eXtreme Gradient Boosting (XGBoost) algorithm. This algorithm optimizes the privacy protection process by applying differential privacy to the optimal segmentation point in a weak classifier. Additionally, to address the multi-party collaboration challenge in federated learning, a differential privacy construction scheme based on multi-party collaboration is proposed. The results indicate that the running times of differential privacy algorithms based on multi-party collaboration, XGBoost, and the traditional differential privacy algorithm were 16.2s, 22.1s, and 29.5s, respectively. The optimized approach improved efficiency by 45.08% compared to the traditional algorithm. Overall, the differential privacy-based federated learning efficiency optimization algorithm can ensure privacy protection while enhancing accuracy and efficiency, providing significant technical support. Introduction: This paper proposes a privacy-preserving joint learning efficiency optimization algorithm based on differential privacy, and designs a differential privacy-preserving algorithm based on XGBoost (DP-XGB). This algorithm enhances privacy preservation by introducing differential privacy at the optimal segmentation point in the weak learner, thereby improving both data security and model accuracy. Building on this foundation, the research further proposes a differential privacy construction scheme (FDP-XGB) based on multi-party collaboration, integrating joint learning techniques to address potential privacy leakage during multi-party collaboration. Objectives: By applying differential privacy to the optimal splitting point among weak learners, DP-XGB optimizes the privacy protection process, thereby enhancing both data security and model accuracy. FDP-XGB is introduced to safeguard privacy in a joint learning environment, effectively addressing the issue of privacy leakage that can occur during multi-party collaboration. Methods: We first enhance the original data and obtains weak learners using the XGBoost algorithm. These weak learners are then combined to form a strong learner, and a differential privacy protection algorithm is constructed. Building on this foundation, the second section develops a multi-party collaborative privacy protection algorithm within a federated learning environment. Results: The results indicate that the running times of differential privacy algorithms based on multi-party collaboration, XGBoost, and the traditional differential privacy algorithm were 16.2s, 22.1s, and 29.5s, respectively. The optimized approach improved efficiency by 45.08% compared to the traditional algorithm. Overall, the differential privacy-based federated learning efficiency optimization algorithm can ensure privacy protection while enhancing accuracy and efficiency, providing significant technical support. Conclusions: This study proposes a privacy protection technology that combines the XGBoost differential privacy protection algorithm with federated learning to address privacy security and data silos in data mining. FDP-XGB demonstrated the highest prediction accuracy when comparing true and predicted data values, outperforming DP-XGB. For a data volume of 18×104, the computation times for XGBoost, DP-XGB, and FDP-XGB were 29.5 seconds, 22.1 seconds, and 16.5 seconds, respectively, with resource consumption rates of 48.5%, 24.9%, and 21.1%.
- Research Article
21
- 10.1109/tdsc.2023.3234599
- Nov 1, 2023
- IEEE Transactions on Dependable and Secure Computing
Federated learning (FL) empowers distributed clients to collaboratively train a shared machine learning model through exchanging parameter information. Despite the fact that FL can protect clients' raw data, malicious users can still crack original data with disclosed parameters. To amend this flaw, differential privacy (DP) is incorporated into FL clients to disturb original parameters, which however can significantly impair the accuracy of the trained model. In this work, we study an imperative question which has been vastly overlooked by existing works: what are the optimal numbers of queries and replies in FL with DP so that the final model accuracy is maximized. In FL, the parameter server (PS) needs to query participating clients for multiple global iterations to complete training. Each client responds a query from the PS by conducting a local iteration. We consider FL that will uniformly and randomly select participating clients to conduct local iterations with the FedSGD algorithm. Our work investigates how many times the PS should query clients and how many times each client should reply the PS by incorporating two most extensively used DP mechanisms (i.e., the Laplace mechanism and Gaussian mechanisms). Through conducting convergence rate analysis, we can determine the optimal numbers of queries and replies in FL with DP so that the final model accuracy can be maximized. Finally, extensive experiments are conducted with publicly available datasets: MNIST and FEMNIST, to verify our analysis and the results demonstrate that properly setting the numbers of queries and replies can significantly improve the final model accuracy in FL with DP.
- Research Article
12
- 10.1016/j.jisa.2022.103309
- Sep 1, 2022
- Journal of Information Security and Applications
High-accuracy low-cost privacy-preserving federated learning in IoT systems via adaptive perturbation
- Research Article
- 10.1002/cpe.70432
- Nov 10, 2025
- Concurrency and Computation: Practice and Experience
The healthcare industry, particularly with the advent of the Internet of Medical Things (IoMT), has witnessed significant integration of Internet of Things (IoT) technologies. IoMT is transforming healthcare by providing substantial benefits to both consumers and healthcare providers. However, the exponential growth in IoMT devices and their data generation raises critical challenges related to data analysis, security, and privacy. Traditional centralized artificial intelligence (AI) approaches, reliant on deep learning (DL) and machine learning (ML) algorithms, struggle to address the increasing complexity of sensitive medical data due to scalability and privacy concerns. Federated Learning (FL) emerges as a promising solution, enabling collaborative model training directly on IoMT devices while preserving data privacy by transmitting only model updates to central servers. This approach ensures data confidentiality and addresses privacy concerns associated with centralized systems. Despite its potential, research on FL in the context of IoMT remains limited. This paper examines the latest developments and innovations in FL, focusing on its application in IoMT and smart healthcare systems. It explores FL architectures, aggregation algorithms, frameworks, and their integration into IoMT‐driven healthcare applications. Additionally, the paper highlights challenges, including data heterogeneity, communication overhead, and security vulnerabilities, alongside privacy‐preserving techniques such as differential privacy, homomorphic encryption (HE), and secure multiparty computation (SMC). Finally, it identifies future research directions to advance FL‐powered IoMT solutions, offering valuable insights for academia and industry stakeholders aiming to enhance privacy‐preserving, intelligent healthcare systems.