• Starts: 1:00 pm on Friday, June 26, 2026
  • Ends: 3:00 pm on Friday, June 26, 2026

ECE PhD Prospectus Defense: Sumatra Dhimoyee

Title: Accelerating LLM Serving Through Storage-Optimized KV-Cache Management

Presenter: Sumatra Dhimoyee

Advisor: Professor Orran Krieger

Chair: Professor Martin Herbordt

Committee: Professor Orran Krieger, Professor Martin Herbordt, Professor Jonathan Appavoo, Dr. Ata Turk

Google Scholar Link: https://scholar.google.com/citations?user=keq4mlUAAAAJ&hl=en

Abstract: Efficient management of the key–value (KV) cache is among the most consequential and difficult problems in large language model (LLM) inference: effective handling can eliminate more than half of redundant GPU computation, cutting both latency and the cost of serving. This proposal develops a multi-tier KV-cache management system that replaces expensive recomputation with low-cost memory and storage access. The massive reuse difference between the prefill and decode phases, and among request types (single- vs. multi-turn, API vs. user-driven) leads the system to leverage offline trace profiling and a global cache directory to control placement, movement and prefetching across a hierarchy of cluster-wide GPU HBM, host DRAM, NVMe and object storage, with all transfers over RDMA. Servers are not statically bound to prefill or decode; requests are mapped dynamically by KV locality and residency. Together these techniques aim to reduce time-to-first-token (TTFT) and inter-token latency (TBT) and to raise goodput at fixed hardware.

Location:
CDS 1001