Abstract
General matrix-matrix multiplication (GeMM) is a principal computational bottleneck in modern deep learning workloads where performance can be heavily impeded by frequent memory accesses. This thesis introduces a multi-dataflow aware processing element (PE) designed to perform the multiply-accumulate (MAC) operations fundamental to GeMMs while utilizing a data flow that best fits the data provided. A salient feature of this architecture is its programmable flexibility, which enables the implementation and execution of different data flow strategies, including Output Stationary, Weight Stationary, Input Stationary, and Row Stationary. The flexible design supports both dense and sparse GeMM and General matrix-vector multiplication (GeMV). To fully utilize the flexibility of the architecture, the objective is to determine the most efficient data flow configurations under various operational constraints, thereby providing critical design insights for future machine learning accelerators.
Publication Date
5-2026
Document Type
Thesis
Student Type
Graduate
Degree Name
Computer Engineering (MS)
Department, Program, or Center
Computer Engineering
College
Kate Gleason College of Engineering
Advisor
Sathwika Bavikadi
Advisor/Committee Member
Cory Merkel
Advisor/Committee Member
Marcin Lukowiak
Recommended Citation
Laudico, Daniel, "Reconfigurable Dataflow for Efficient Matrix-Matrix Multiplication for DNN Acceleration" (2026). Thesis. Rochester Institute of Technology. Accessed from
https://repository.rit.edu/theses/12780
Campus
RIT – Main Campus
