Abstract

Data-intensive workloads in high-performance computing and artificial intelligence move massive volumes of data across deep memory and storage hierarchies, often spending more time moving data than computing on it. Their access patterns, data-sharing relationships, and computational graph properties are largely determined by descriptors known to the application before execution. However, general-purpose runtimes treat this structure as opaque, and the resulting gap between workload behavior and runtime decisions accounts for a substantial fraction of their execution time. Despite advances in performance modeling and data-movement systems, this structure remains underexploited. Schedulers, caches, and I/O paths provision resources conservatively, prediction models are rarely reusable across workloads, and regularities such as shared datasets, spatial clustering of scattered accesses, and overlap across requests are not often exposed to the runtime. This dissertation develops workload-aware techniques that make runtimes aware of this structure at different execution stages. At provisioning time, we propose a runtime that configures deep-learning training containers from the model class, dataset, and batch size, targeting serverless deployment where rigid resource limits make provisioning especially hard, reducing training time by 46% and memory use by 44% on average. For scheduling purposes, we introduce a predictor that estimates distributed training time from a reusable embedding of the computational graph, with up to 9.8× lower error than sampling-based baselines, and a technique that staggers co-located instances to maximize cache reuse, cutting their resource utilization by more than 50%. As workloads read and reuse data, we propose a runtime that coalesces the scattered small reads dominating large-scale analytics into larger operations tuned against downstream computation, with an average of 3.6× higher throughput than an asynchronous I/O baseline, and an indexing mechanism that adapts to the request-length distribution during large language model inference, reducing time-to-first-token by up to 3.5×. To reinforce scientific integrity, we design a checkpoint-based reproducibility analytics framework that localizes divergence via tree-based metadata comparison rather than element-wise, up to 11× faster than direct comparison. These contributions advance workload-aware runtimes from performance prediction to data-movement coordination across scientific simulation, deep-learning training, and language model inference. They show that access-pattern regularities and data-sharing relationships routinely ignored by general-purpose runtimes carry enough information to deliver substantial gains at low overhead.

Publication Date

9-2026

Document Type

Dissertation

Student Type

Graduate

Degree Name

Computing and Information Sciences (Ph.D.)

Department, Program, or Center

Computing and Information Sciences Ph.D, Department of

College

Golisano College of Computing and Information Sciences

Advisor

M. Mustafa Rafique

Advisor/Committee Member

Minseok Kwon

Advisor/Committee Member

Bogdan Nicolae

Campus

RIT – Main Campus

Share

COinS