Abstract

With advancement in technology, the needs for high performance computing are increasing tremendously. Cluster computing has developed due to the availability of high performance cost effective processors and high speed networks. The long-term trend in High performance computing requires increasing number of nodes in parallel computing platforms. This however entails a higher failure probability. The Message Passing Paradigm (MPI) is currently the programming paradigm and communication library most commonly used on parallel computing platforms. MPI applications may get stopped at any time due to unpredictable failures during execution. In our paper we propose an efficient fault tolerant approach for MPI system in an asymmetric cluster computing environment. In this paper, we use centralized logging process. In the approach proposed, we use message logging for message losses. The process has three main parts failure detection, failure recovery and overload detection. Our System maintains monitor nodes for all nodes in cluster, the difference being all monitor nodes can work as a cluster node even when the system is functioning properly and not just at the time of node failure.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call