- Starts: 1:00 pm on Friday, June 26, 2026
- Ends: 3:00 pm on Friday, June 26, 2026
ECE PhD Prospectus Defense: Sumatra Dhimoyee
Title: Accelerating LLM Serving Through Storage-Optimized KV-Cache Management
Presenter: Sumatra Dhimoyee
Advisor: Professor Orran Krieger
Chair: Professor Martin Herbordt
Committee: Professor Orran Krieger, Professor Martin Herbordt, Professor Jonathan Appavoo, Dr. Ata Turk
Google Scholar Link: https://scholar.google.com/citations?user=keq4mlUAAAAJ&hl=en
Abstract: Efficient management of the key–value (KV) cache is among the most consequential and difficult problems in large language model (LLM) inference: effective handling can eliminate more than half of redundant GPU computation, cutting both latency and the cost of serving. This proposal develops a multi-tier KV-cache management system that replaces expensive recomputation with low-cost memory and storage access. The massive reuse difference between the prefill and decode phases, and among request types (single- vs. multi-turn, API vs. user-driven) leads the system to leverage offline trace profiling and a global cache directory to control placement, movement and prefetching across a hierarchy of cluster-wide GPU HBM, host DRAM, NVMe and object storage, with all transfers over RDMA. Servers are not statically bound to prefill or decode; requests are mapped dynamically by KV locality and residency. Together these techniques aim to reduce time-to-first-token (TTFT) and inter-token latency (TBT) and to raise goodput at fixed hardware.
- Location:
- CDS 1001
