Conferences >2014 Hardware-Software Co-Des...

An Implementation of Block Conjugate Gradient Algorithm on CPU-GPU Processors

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

In this paper, we investigate the implementation of the Block Conjugate Gradient (BCG) algorithm on CPU-GPU processors. By analyzing the performance of various matrix ope...Show More

Metadata

Abstract:

In this paper, we investigate the implementation of the Block Conjugate Gradient (BCG) algorithm on CPU-GPU processors. By analyzing the performance of various matrix operations in BCG, we identify the main performance bottleneck in constructing new search direction matrices. Replacing the QR decomposition by eigendecomposition of a small matrix remedies the problem by reducing the computational cost of generating orthogonal search directions. Moreover, a hybrid (offload) computing scheme is designed to enables the BCG implementation to handle linear systems with large, sparse coefficient matrices that cannot fit in the GPU memory. The hybrid scheme offloads matrix operations to GPU processors while helps hide the CPU-GPU memory transaction overhead. We compare the performance of our BCG implementation with the one on CPU with Intel Xeon Phi coprocessors using the automatic offload mode. With sufficient number of right hand sides, the CPU-GPU implementation of BCG can reach speedup of 2.61 over the CPU-only implementation, which is significantly higher than that of the CPU-Intel Xeon Phi implementation.

Published in: 2014 Hardware-Software Co-Design for High Performance Computing

Date of Conference: 17-17 November 2014

Date Added to IEEE Xplore: 22 January 2015

Electronic ISBN:978-1-4799-7564-8

DOI: 10.1109/Co-HPC.2014.10

Conference Location: New Orleans, LA, USA

Contents

References is not available for this document.

An Implementation of Block Conjugate Gradient Algorithm on CPU-GPU Processors

Abstract:

Metadata

Abstract:

References

IEEE Account

Purchase Details

Profile Information

Need Help?

An Implementation of Block Conjugate Gradient Algorithm on CPU-GPU Processors

Alerts

Abstract:

Metadata

Abstract:

Authors

Figures

References

Citations

Keywords

Metrics

References

IEEE Account

Purchase Details

Profile Information

Need Help?