Edge–Cloud Collaboration for Machine Condition Monitoring: A Comprehensive Review of Mechanisms, Models, and Applications
This review of 147 publications from 2019 to 2026 explores edge–cloud collaboration for machine condition monitoring, categorizing mechanisms into four dimensions and highlighting technologies like federated and transfer learning. It finds model-centric approaches dominate, with emerging focus on architecture and trust, and identifies key challenges such as generalization, data efficiency, interoperability, and trustworthy deployment, outlining future research directions including foundation models and privacy-preserving systems.
Machine condition monitoring increasingly depends on distributed sensing, edge intelligence, and cloud analytics, yet timely and trustworthy health assessment remains constrained by latency, bandwidth, privacy, and reliability requirements. Cloud-only architectures provide scalable computation and historical data integration but often fail to satisfy real-time industrial needs, whereas edge-only deployments are limited by restricted computing resources and fragmented local knowledge. Edge–cloud collaboration has, therefore, emerged as a practical architecture for distributing perception, inference, learning, and coordination across hierarchical industrial systems. This review examines 147 publications on edge–cloud collaboration for machine condition monitoring published between 2019 and February 2026. A four-dimensional taxonomy is developed to organize the literature into model-centric, data-centric, resource and task-centric, and architecture and trust-centric mechanisms, while 13 survey and review papers are considered separately for contextual comparison. On this basis, the review analyzes representative collaboration mechanisms and enabling technologies, with particular attention to federated learning, transfer learning, knowledge distillation, digital twins, and deep reinforcement learning, and surveys their deployment in manufacturing, energy, transportation, and infrastructure monitoring scenarios. The literature remains dominated by model-centric collaboration, while architecture and trust-centric studies increasingly provide the system foundations required for practical deployment. The review further identifies major open challenges, including robust generalization under changing operating conditions, efficient data transmission, real-time resource coordination, interoperability, and trustworthy large-scale deployment, and outlines future directions in foundation-model-based edge–cloud collaboration, continual learning, dual digital twins, trustworthy collaboration, and privacy-preserving industrial ecosystems.
- Research Article
- 10.23919/jcc.2023.10061658
- Feb 1, 2023
- China Communications
With the growing maturity of the advanced edge-cloud collaboration and integrated sensing-communication-computing systems, edge intelligence has been envisioned as one of the enabling technologies for ubiquitous and latency-sensitive machine learning based services in future wireless systems. However, with the ever-growing scale of the edge-cloud systems as well as the rapid advances in future 6G technologies, there exists a critical issue that limits the deep penetration of future 6G enabled edge intelligence services, i.e., how to accurately evaluate the performance when leveraging the emerging yet pre-matured technologies, protocols, and algorithms into the existing systems. The recent advanced Digital Twin (DT) has been envisioned to provide a fault-tolerant and low-cost platform for accurately simulating and evaluating the emerging technologies, protocols, and algorithms, without causing any negative influence on the existing systems. The essence of DT aims at enabling a systematic and fully digitalized modeling systems that can produce accurate digital model of the corresponding physical identity in the DT-space. By simulating the digital models in the DT-space, DT can accurately and reliably predict and estimate the dynamics and evolutions of the physical networks during the entire life cycle, thus providing a timely and risk-free methodology for evaluating the performance when adopting the emerging yet pre-matured technologies, protocols, and algorithms for network management. Therefore, DT has been considered as an efficient scheme for addressing the challenging issue to 6G enabled edge intelligence, and thus has attracted lots of interests from both academia and industries. It is observed that DT and edge intelligence can benefit each other, which thus yields a deep convergence of themselves. On the one hand, the increasing network scale, complexity, and security risks of edge intelligence raise more challenges to the security and reliability. DT can accurately simulate and predict the performance of edge intelligence services and thus provide accurate benchmark references for edge services. On the other hand, to enable accurate mapping and real-time synchronization between the physical identifies and their digital models, DT necessitates massive sensing of the targeted physical identifies as well as the consequent data analytics and modeling. Moreover, efficient yet low-latency transmissions of the sensing data and the data analytics (e.g., the inference models) are required. 6G edge intelligence can naturally provide a solution to these requirements. Therefore, this special issue focuses on the convergence of DT and 6G enabled Edge Intelligence, from the perspectives of theories, algorithms, and applications.
- Conference Article
27
- 10.1109/tst52996.2021.00019
- May 1, 2021
Throughout recent years, many "Smart Factory" concepts have been emerged to escort the current technological progress. Among these concepts, we found the Digital Twin which is a burgeoning technology that attracts great interests from academics and industry. Akin to other technologies, the power of Digital Twin can be strengthened by leveraging other technologies' benefits such as those of cloud computing and edge computing. First, to support the implementation of the Digital Twin for the shop floor monitoring, a conceptual architecture of the Digital Twin based on edge-cloud collaboration is proposed. This architecture is believed to enhance the real-time capabilities of the Digital Twin while dealing with the abundant amounts of manufacturing data. Moreover, the paper discusses how the microservices architecture can benefit the proposed Digital Twin architecture. On the other hand, in order to provide an underpinning to researchers on the Digital Twin topic, the paper sums up the description and an up-to-date output of the Digital Twin, highlighting four manufacturing application scenarios.
- Conference Article
27
- 10.1109/indin45523.2021.9557492
- Jul 21, 2021
The physical and virtual access of manufacturing equipment to the cloud platform is an important direction of cloud manufacturing research, but the current research results in this direction are not enough to support the "landing" and wide application of cloud manufacturing model. In the traditional cloud manufacturing model, the direct access of manufacturing equipment to the cloud brings huge network pressure and unbearable latency, and also poses a privacy and security risk. Aiming at the cloudification of manufacturing equipment in the cloud manufacturing environment, the 3D printer is selected as the research object, and a cloud-edge collaboration architecture for 3D printers based on digital twin is proposed. To solve the lack of network reliability brought about by traditional cloudification, time-sensitive services are deployed at the edge, realizing the local extension of cloud services. The digital twin information model of FDM additive manufacturing is defined. Real time controlling and monitoring based on digital twin is realized. By cloud-edge collaboration, an application case is implemented, which verifies the effectiveness of the system and provides a reference solution for the cloudification of other manufacturing equipment in the environment of cloud manufacturing.
- Research Article
- 10.31449/inf.v50i7.9320
- Feb 21, 2026
- Informatica
This study proposes a novel adaptive control framework integrating nonlinear IoT dynamics with mechatronic systems to address the challenges of strong coupling, uncertainty, and real-time constraints. Our key innovation lies in a hierarchical optimization architecture combining model predictive control (MPC) with deep reinforcement learning (DRL), enabling dynamic adaptation through edge-cloud collaboration. Experimental results demonstrate efficiency improvements from 13.79% to 83.46% in subsystem performance, highlighting the algorithm's capability to balance robustness and adaptability. This work fills a critical gap in existing methods by unifying distributed sensing, online learning, and nonlinear optimization for IoT-enabled mechatronic systems. In contrast, the presence of high efficiency values, such as 93.76% and 80.75%, in certain parts of the system indicates that the system is able to achieve efficient operation under certain specific conditions. This study proposes a novel adaptive control framework integrating nonlinear IoT dynamics with mechatronic systems, leveraging Model Predictive Control (MPC), Deep Reinforcement Learning (DRL), and Differential Evolution (DE) within a hierarchical optimization architecture. Experimental validation on a manipulator testbed demonstrates 83.46% efficiency improvement in subsystem coordination, <100 ms response time under dynamic coupling, and 92.5% trajectory accuracy with ±5% standard deviation under 30% noise interference.
- Research Article
60
- 10.1109/tii.2021.3113875
- Jun 1, 2022
- IEEE Transactions on Industrial Informatics
The industrial Internet of Things (IIoT) and 5G have been served as the key elements to support the reliable and efficient operation of Industry 4.0. By integrating burgeoning network function virtualization (NFV) technology with cloud computing and mobile edge computing, an NFV-enabled cloud–edge collaborative IIoT architecture can efficiently provide flexible service for the massive IIoT traffic in the form of a service function chain (SFC). However, the efficient cloud–edge collaboration, the reasonable comprehensive resource consumption, and different quality of services are still key problems to be solved. Thus, to balance the quality of IIoT services, as well as computational and communicational resource consumption, a multiobjective SFC deployment model is designed to characterize the diverse service requirements and specific network environment for the IIoT. Then, a deep- <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$Q$</tex-math></inline-formula> -learning-based online SFC deployment algorithm is presented, which can efficiently learn the relationship between the SFC deployment scheme and its performance through the iterative training. Simulation results demonstrate that our proposed approach outperforms others in balancing the resource consumption, accepting more SFC requests, as well as providing differentiated services for delay-sensitive IIoT traffic and resource-intensive IIoT traffic.
- Research Article
1
- 10.1109/tce.2024.3445916
- May 1, 2025
- IEEE Transactions on Consumer Electronics
The vehicle-road information interaction based on cloud-edge collaboration is an important part of the future intelligent transportation system. By interacting with the integrated data shared to the cloud by edge nodes and roadside units (RSU), vehicles as information consumers can perceive road information in real-time, thereby reducing traffic accidents and ensuring driving safety. However, challenges such as long communication latency between the cloud and vehicles, limited resources at the edge, and high mobility of vehicles lead to problems such as discontinuous information interaction and long synchronization latency in complex traffic scenarios. To address the above problems, we propose a low-latency vehicle information synchronization scheme. The scheme relies on digital twins to map real-time traffic scenarios to ensure information continuity. It allocates computational resources through sequential least squares to reduce the synchronization update latency. Since the limited coverage of edge nodes leads to high latency of interactions, we develop an update and migration optimization algorithm based on deep reinforcement learning and reduce the average total vehicle latency by restricting each migration decision to a local pre-selection of edge nodes. Based on extensive experimentation with real-world vehicle movement datasets, our approach can reduce the latency by 50% compared to existing baseline methods.
- Research Article
27
- 10.1016/j.asoc.2023.110082
- Feb 10, 2023
- Applied Soft Computing
Innovative soft computing-enabled cloud optimization for next-generation IoT in digital twins
- Research Article
- 10.4018/ijitsa.337797
- Feb 7, 2024
- International Journal of Information Technologies and Systems Approach
In order to solve the problems that most models are complex, time-consuming, and have difficulty in identifying image errors, an image identification and error correction method of test report based on deep reinforcement learning and the internet of things platform in the smart lab was proposed. Firstly, a smart lab architecture was designed based on the internet of things platform, achieving efficient operation of the laboratory through cloud edge collaboration. Then, the depth separable convolution improved convolutional neural network is used to extract image features, and the features are input into bidirectional recurrent neural networks (BiLSTM) for analysis to complete image recognition. Finally, the ICNN-BiLSTM model is used as the agent of reinforcement learning, and image error correction is completed by identifying the distance between the image and the key points of the reference image. Based on the Python platform, the proposed method was experimentally demonstrated, and the results showed that its average error correction accuracy reached 96.75%, with a processing time of 15.37s.
- Research Article
- 10.3390/math13193055
- Sep 23, 2025
- Mathematics
Current cloud–edge collaboration collaboration architectures face challenges in security resource scheduling due to their mostly static nature, which cannot keep up with real-time attack patterns and dynamic security needs. To address this, this paper proposes a dynamic scheduling method using Deep Reinforcement Learning (DQN) and SRv6 technology. The method establishes a multi-dimensional feature space by collecting network threat indicators and security resource states; constructs a dynamic decision-making model with DQN to optimize scheduling strategies online by encoding security requirements, resource constraints, and network topology into a Markov Decision Process; and enables flexible security service chaining through SRv6 for precise policy implementation. Experimental results demonstrate that this approach significantly reduces security service deployment delays (by up to 56.8%), enhances resource utilization, and effectively balances the security load between edge and cloud.
- Research Article
7
- 10.1109/tce.2024.3454270
- Jan 1, 2024
- IEEE Transactions on Consumer Electronics
The rapid development of vehicular networks has led to widespread adoption of various electric vehicle (EV) applications, often necessitating the deployment of large-scale deep neural networks (DNNs). However, constrained computational and energy resources pose challenges for executing computationally intensive DNN tasks exclusively within EVs. To address this issue, one potential solution is to utilize edge or cloud computing resources for collaborative computation, typically implemented through DNN partitioning and task offloading. Therefore, we propose a novel approach in this paper, named TOCC, to execute EV-generated DNN tasks with edge-cloud collaboration. Specifically, we first construct a performance prediction model, which can accurately predict the performance of different layers in a DNN. Then, we define the joint optimization problem of minimizing processing delay and energy consumption as a Markov Decision Process (MDP). Finally, we employ Deep Reinforcement Learning (DRL) to design a strategy that enables EVs to make optimal decisions for DNN task partitioning and offloading. The experimental results demonstrate that TOCC surpasses existing approaches in terms of both processing delay and energy consumption, and is applicable to various DNN types.
- Research Article
4
- 10.1109/jiot.2025.3550916
- Jun 15, 2025
- IEEE Internet of Things Journal
In intelligent manufacturing systems, accurate and timely fault diagnosis is crucial for ensuring a safe and stable manufacturing process. While transfer learning (TL) can mitigate the need for extensive labeled data, not all historical datasets are applicable to specific fault diagnosis tasks, and the use of inappropriate datasets can deteriorate the accuracy of TL models. To address these issues, a TL fault diagnosis method based on cloud-edge collaboration is proposed. First, a variational autoencoder TL algorithm based on multiscale convolution and domain fusion (MSDF-VAE) is presented to effectively leverage extensive historical fault data, particularly in scenarios with limited labeled samples. Second, a lightweight autoencoder model (LAE) is employed to improve the reusability and specificity of historical data and fault diagnosis models by analyzing the correlation between historical data and the current data. Additionally, to reduce latency and meet real-time requirements for fault diagnosis tasks, a cloud-edge collaborative framework is proposed, within which MSDF-VAE and LAE are deployed. This approach enables real-time diagnosis using the MSDF-VAE model at the edge layer, while the cloud layer concurrently trains a high-precision model with the selected data by the LAE. The experiments verify the accuracy of the MSDF-VAE and confirm the effectiveness of the proposed cloud-edge collaboration framework.
- Conference Article
5
- 10.1115/msec2020-8355
- Sep 3, 2020
- Volume 2: Manufacturing Processes; Manufacturing Systems; Nano/Micro/Meso Manufacturing; Quality and Reliability
Cloud manufacturing is a service-oriented networked manufacturing model that embraces the concept of ‘Everything-as-a-Service’. In cloud manufacturing, distributed manufacturing resources encompassed in the product lifecycle are transformed into manufacturing services. Industrial robots are an important category of manufacturing resources in cloud manufacturing. During the past years, robots have been demonstrated to be able to learn various dexterous manipulation skills through training with deep reinforcement learning (DRL). In cloud manufacturing, there are many complex industrial application scenarios that require dexterous robots. Hence, robot training, which enables robots to learn various manipulation skills, becomes an important requirement for cloud manufacturing in the future, leading to the concept of ‘Robot Training-as-a-Service’. This paper focuses on industrial robot training in the context of cloud manufacturing. First, related work on cloud manufacturing, DRL, DRL-based robot training, and cloud-edge collaboration is briefly reviewed and analyzed. Then, a framework for industrial robot training in cloud manufacturing with DRL is proposed, and a simplified case study is presented to demonstrate the basic principle of the framework. Finally, possible future research issues are discussed.
- Book Chapter
- 10.71443/9789349552647-09
- Mar 18, 2026
Next-generation wireless communication systems require fast, reliable, and efficient data transmission to support modern applications such as smart cities, autonomous systems, and IoT networks. Antenna arrays play an important role in improving signal quality and coverage through beamforming and spatial communication. Machine learning techniques help in making communication systems more intelligent by enabling data-driven optimization of processes such as channel estimation, beam selection, and interference management. The combination of antenna arrays and machine learning supports the development of intelligent edge communication, where data processing happens close to the source, reducing delay and improving system performance. This integration allows real-time decision-making, better resource utilization, and improved network efficiency. Edge–cloud collaboration further enhances system capabilities by balancing computational tasks between local devices and centralized servers. This chapter discusses the importance of integrating antenna arrays with machine learning, focusing on system optimization, edge intelligence, and multimodal data fusion. It also highlights key challenges such as synchronization issues, computational limitations, and data requirements. The overall approach contributes to the development of adaptive and efficient wireless communication systems for future networks.
- Conference Article
1
- 10.1117/12.2662510
- Dec 29, 2022
Based on the needs of building real traffic twin scenes, this paper proposes a traffic twin scene scheme architecture based on edge-cloud collaboration, and introduces the key technologies involved. This paper explores the data processing scheme of multi-node cooperation when the road traffic flow is large, and combines the detection and identification technology of traffic participants to complete the accurate mapping of traffic participants and twin models, and establish the driving and exit mechanism of the real traffic participant model to achieve High-fidelity digital traffic twin scenarios.
- Research Article
246
- 10.1109/jiot.2021.3098508
- Nov 15, 2021
- IEEE Internet of Things Journal
Sixth-generation (6G) is envisioned to be characterized by ubiquitous connectivity, extremely low latency, and enhanced edge intelligence. However, enriching 6G with these features requires addressing new, unique, and complex challenges specifically at the edge of the network. In this article, we propose a wireless digital twin edge network model by integrating digital twin with edge networks to enable new functionalities, such as hyper-connected experience and low-latency edge computing. To efficiently construct and maintain digital twins in the wireless digital twin network, we formulate the edge association problem with respect to the dynamic network states and varying network topology. Furthermore, according to the different running stages, we decompose the problem into two subproblems, including digital twin placement and digital twin migration. Moreover, we develop a deep reinforcement learning (DRL)-based algorithm to find the optimal solution to the digital twin placement problem, and then use transfer learning to solve the digital twin migration problem. Numerical results show that the proposed scheme provides reduced system cost and enhanced convergence rate for dynamic network states.