ECE PhD Thesis Defense: Efe Sencan
- Starts: 9:45 am on Wednesday, July 22, 2026
- Ends: 11:45 am on Wednesday, July 22, 2026
ECE PhD Thesis Defense: Efe Sencan
Title: Automated Anomaly Detection and Performance Diagnosis for Production HPC Systems
Presenter: Efe Sencan
Advisors: Professors Ayșe Coskun & Brian Kulis
Chair: Professor Ashok Cutkosky
Committee: Professor Ayșe Coskun, Professor Brian Kulis, Professor Manuel Egele, Professor Wenchao Li, Dr. Dhruva Kulkarni
Google Scholar Link: https://scholar.google.com/citations?user=RwrhVIcAAAAJ&hl=tr&oi=ao
Advisor: Modern production high-performance computing (HPC) systems execute diverse scientific workloads across thousands of compute nodes and increasingly rely on heterogeneous, GPU-accelerated architectures. Subtle performance degradations caused by hardware faults, resource contention, inefficient resource allocation, or GPU kernel-level bottlenecks can reduce system throughput and increase operational cost without necessarily causing application failures. Detecting, characterizing, and diagnosing these issues at scale requires automated methods that can operate under the practical constraints of production environments, where labeled data is scarce, training data may be contaminated, and detailed profiling tools can introduce substantial overhead.
In this dissertation, we develop automated methods for performance analysis and diagnosis across multiple levels of large-scale HPC systems, from node-level telemetry to workload-level GPU utilization and kernel-level bottlenecks. We first present Refine, a robust unsupervised anomaly detection framework for contaminated production HPC telemetry. We then introduce metrics for characterizing GPU resource utilization and imbalance across large-scale production workloads. Finally, we develop an automated GPU kernel bottleneck diagnosis framework that uses hardware-counter measurements to identify performance-limiting kernel behaviors and reduce the cost of GPU profiling. These methods reduce the reliance on labeled data and exhaustive profiling, making performance diagnosis more practical for production HPC environments.
- Location:
- PHO 339