Achieving scalable cluster system analysis and management with a gossip-based network service

D.E Collins,R.A Quander,A.D George

doi:10.1109/lcn.2001.990767

Abstract

Clusters of workstations are increasingly used for applications requiring high levels of both performance and reliability. Certain fundamental services are highly desirable to achieve these twin goals of network-based cluster system analysis and management. Among these services is the ability to detect network and node failures and the capability to efficiently determine computer and network load levels. Furthermore, the ability to allow for the distribution of administrative directives is also integral to the goal of cluster management. This paper presents a scalable approach to providing these vital support capabilities for distributed computing integrated into a cluster management system. Previous approaches to cluster management have suffered from problems of scalability and the inability to properly support heterogeneous systems in a non-proprietary fashion. This cluster management system employs gossip techniques to address the problem of scalability in network-based system management. The results of two case studies show that the cluster management system is scalable and has little adverse impact on the performance of sequential and parallel applications running on the managed system.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Achieving scalable cluster system analysis and management with a gossip-based network service

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Single Chip Microcomputer Cluster Management
...
Advanced Materials Research | VOL. 933
, et. al. ...
01 May 2014
Advanced Materials Research | VOL. 933

SIMULATION OF HIERARCHICAL RESOURCE MANAGEMENT FOR META-COMPUTING SYSTEMS
J Santoso ... G D Van Albada
International Journal of Foundations of Computer Science | VOL. 12
J Santoso, et. al.J Santoso ... G D Van Albada
01 Oct 2001
International Journal of Foundations of Computer Science | VOL. 12

Data Distribution Management Modeling and Implementation on Computational Grid
Jong Sik Lee
-
Jong Sik LeeJong Sik Lee
01 Jan 2004
01 Jan 2004

Raptor: integrating checkpoints and thread migration for cluster management
H Shafi ... E Speight
-
H Shafi, et. al.H Shafi ... E Speight
20 Oct 2003
20 Oct 2003

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Achieving scalable cluster system analysis and management with a gossip-based network service

Abstract

Talk to us

Similar Papers