SCHOLARLY PUBLICATION
Anatomy of high-performance matrix multiplication
📄 Abstract
We present the basic principles that underlie the high-performance implementation of the matrix-matrix multiplication that is part of the widely used GotoBLAS library. Design decisions are justified by successively refining a model of architectures with multilevel memories. A simple but effective algorithm for executing this operation results. Implementations on a broad selection of architectures are shown to achieve near-peak performance.
📤 Share this page
Found this useful? Share it with your network.