Abstract

While a strong history of computer development has maximized architecture performance, modern computing is fundamentally constrained by thermal limits and power consumption. This is especially true in power sensitive edge applications, which often have low activity. Asynchronous architectures offer a pathway to take advantage of sparsity within hardware to drastically reduce dynamic power through fine-grain selective activation of only necessary hardware. Despite these advantages, the semiconductor industry relies heavily on Electronic Design Automation (EDA) toolchains, which have been developed with many synchronous assumptions. The traditional overhead and strict incompatibility with standard EDA toolchains have limited adoption of asynchronous technologies in widespread commercial designs. This research hopes to remove these barriers by establishing a comprehensive methodology to design, synthesize, and validate asynchronous architectures using existing commercial toolchains. Through a literature search, bundled-data asynchronous protocols were chosen to develop the following work. An optimized library of asynchronous primitives was developed through a process of literature review and identifying optimized circuits after identifying assumptions made in prior work. This led to a comparison of handshake cell and control-flow circuits to determine the most efficient implementations for this design methodology. The most practical and efficient circuits were characterized to provide an analytical framework for architects to choose the best phase protocols. A set of scripts were developed for local clock tree and matched delay insertion for bundled data architectures following design compilation for pre-layout verification. The practical application of these methodologies, circuits, and tools was proven through the design and characterization of a standalone, asynchronous data-driven cache architecture. Compared to its synchronous counterpart, the asynchronous design achieved a 45.7% decrease in latency with only a marginal 1.73% increase in total area. This parameterized design was characterized under various stimuli to prove its viability, particularly in scenarios with inherent data sparsity. The results showed significant efficiency gains: under low sparsity with a low hit rate, the asynchronous cache used 65.50% more dynamic power on average, while under high sparsity and a higher hit rate, its dynamic power dropped by 84.57% on average compared to the synchronous baseline. When evaluating total power for the 28nm design, static power dominates across most configurations and hit rates. As a result, the asynchronous design generally consumes higher total power, and once DRAM access times exceed 12 ns, the total power of the asynchronous design consistently drops below the synchronous model. Ultimately, this work provides a rigorous and accessible framework to help future engineers develop scalable, power-aware digital systems using existing industry tools.

Publication Date

8-2026

Document Type

Thesis

Student Type

Graduate

Degree Name

Electrical Engineering (MS)

Department, Program, or Center

Electrical Engineering

College

Kate Gleason College of Engineering

Advisor

Mark A. Indovina

Advisor/Committee Member

Dorin Patru

Advisor/Committee Member

Carlos Barrios

Campus

RIT – Main Campus

Share

COinS