By Topic

Troubleshooting thousands of jobs on production grids using data mining techniques

Sign In

Cookies must be enabled to login.After enabling cookies , please use refresh or reload or ctrl+f5 on the browser for the login options.

Formats Non-Member Member
$33 $13
Learn how you can qualify for the best price for this item!
Become an IEEE Member or Subscribe to
IEEE Xplore for exclusive pricing!
close button

puzzle piece

IEEE membership options for an individual and IEEE Xplore subscriptions for an organization offer the most affordable access to essential journal articles, conference papers, standards, eBooks, and eLearning courses.

Learn more about:

IEEE membership

IEEE Xplore subscriptions

3 Author(s)
David A. Cieslak ; Department of Computer Science and Engineering, University of Notre Dame, USA ; Nitesh V. Chawla ; Douglas L. Thain

Large scale production computing grids introduce new challenges in debugging and troubleshooting. A user that submits a workload consisting of tens of thousands of jobs to a grid of thousands of processors has a good chance of receiving thousands of error messages as a result. How can one begin to reason about such problems? We propose that data mining techniques can be employed to classify failures according to the properties of the jobs and machines involved. We demonstrate this technique through several case studies on real workloads consisting of tens of thousands of jobs. We apply the same techniques to a yearpsilas worth of data on a 3000 CPU production grid and use it to gain a high level understanding of the system behavior.

Published in:

2008 9th IEEE/ACM International Conference on Grid Computing

Date of Conference:

Sept. 29 2008-Oct. 1 2008