Algorithm-based recovery for HPL

Teresa Davies,Christer Karlsson,Hui Liu,Zizhong Chen

doi:10.1145/2038037.1941600

Algorithm-based recovery for HPL

Teresa Davies, Christer Karlsson + Show 2 more

https://doi.org/10.1145/2038037.1941600

Copy DOI

Journal: ACM SIGPLAN Notices

Publication Date: Feb 12, 2011

Affiliation: Colorado School of Mines

#Diskless Checkpointing #Large Amount Of Data + Show 8 more

Abstract
Full-Text
Similar Papers

Abstract

When more processors are used for a calculation, the probability that one will fail during the calculation increases. Fault tolerance is a technique for allowing a calculation to survive a failure, and includes recovering lost data. A common method of recovery is diskless checkpointing. However, it has high overhead when a large amount of data is involved, as is the case with matrix operations. A checksum-based method allows fault tolerance of matrix operations with lower overhead. This technique is applicable to the LU decomposition in the benchmark HPL.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: ACM SIGPLAN Notices

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.