Masters Theses

Keywords and Phrases

data transfer optimization; distributed systems; high-performance networks; reinforcement learning

Abstract

"The end-to-end performance of scientific data movement is increasingly limited not by raw network capacity but by whichever resource along the path is momentarily slowest, a bottleneck that no single static configuration can track. This thesis argues that the right response is to decompose each problem along the resource or workflow axis where bottlenecks vary, adapt only the genuinely dynamic control variables at runtime, and make that adaptation efficient through utility-guided or offline-learned control rather than expensive online search. The thesis develops this principle through four systems of widening scope. AutoMDT decomposes a transfer by operation, assigning separate concurrency levels to the read, network, and write stages, and training a reinforcement learning agent in an offline simulator. LDM decomposes by workload structure, assigning pipelining and parallelism structurally and adapting only chunk-level concurrency, so heterogeneous file-size distributions are handled without an exponential online search. FastBioDL moves the principle to public genomic repositories, where users control only the client, performing adaptive segmented retrieval tuned by a utility-guided controller. SeqFlux lifts the target to the full genomic acquisition pipeline, overlapping download, conversion, and compression under telemetry-driven admission control and phase-aware disk reservation. Across production grade testbeds, the four systems reduce completion time, sustain higher throughput on heterogeneous and constrained paths, and accelerate end-to-end acquisition while using fewer resources than aggressive static baselines. Together they show that scientific data movement is better served by structured decomposition and efficient online adaptation than by static tuning or exhaustive search"-- Abstract, p. iv

Advisor(s)

Arifuzzaman, Md

Committee Member(s)

Imran, Mia Mohammad
Puri, Satish

Department(s)

Computer Science

Degree Name

M.S. in Computer Science

Publisher

Missouri University of Science and Technology

Publication Date

2026

Journal article titles appearing in thesis/dissertation

Paper I: Pages 11–38 have been published as R. M. Swargo, E. Arslan, and M. Arifuzzaman, “Modular Architecture for High-Performance and Low Overhead Data Transfers,” in Proceedings of the SC ’25Workshops of the International Conference for High Performance Computing, Networking, Storage, and Analysis (INDIS), ACM, 2025, pp. 939– 948.

Paper II: Pages 39–71 have been submitted as R. M. Swargo and M. Arifuzzaman, “Unified Multi-Parameter Data Transfer Optimization for Heterogeneous Workloads,” to the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC ’26).

Paper III: Pages 72–98 have been submitted as R. M. Swargo, N. H. Neom, and M. Arifuzzaman, “Self-Tuning Segmented Retrieval of Large-Scale Genomic Data,” to the IEEE International Conference on e-Science (eScience ’26).

Paper IV: Pages 99–143 have been submitted as R. M. Swargo, N. H. Neom, and M. Arifuzzaman, “SeqFlux: Resource-Aware Pipeline Coordination for Genomic Data Acquisition on HPC Clusters,” to the Future Generation Computer Systems (FGCS) journal.

Pagination

xvi, 155 pages

Note about bibliography

Includes_bibliographical_references_(pages 148-153)

Rights

© 2026 Rasman Mubtasim Swargo , All Rights Reserved

Document Type

Thesis - Open Access

File Type

text

Language

English

Thesis Number

T 12628

Share

 
COinS