Abstract

General matrix-matrix multiplication (GeMM) is a principal computational bottleneck in modern deep learning workloads where performance can be heavily impeded by frequent memory accesses. This thesis introduces a multi-dataflow aware processing element (PE) designed to perform the multiply-accumulate (MAC) operations fundamental to GeMMs while utilizing a data flow that best fits the data provided. A salient feature of this architecture is its programmable flexibility, which enables the implementation and execution of different data flow strategies, including Output Stationary, Weight Stationary, Input Stationary, and Row Stationary. The flexible design supports both dense and sparse GeMM and General matrix-vector multiplication (GeMV). To fully utilize the flexibility of the architecture, the objective is to determine the most efficient data flow configurations under various operational constraints, thereby providing critical design insights for future machine learning accelerators.

Publication Date

5-2026

Document Type

Thesis

Student Type

Graduate

Degree Name

Computer Engineering (MS)

Department, Program, or Center

Computer Engineering

College

Kate Gleason College of Engineering

Advisor

Sathwika Bavikadi

Advisor/Committee Member

Cory Merkel

Advisor/Committee Member

Marcin Lukowiak

Campus

RIT – Main Campus

Share

COinS