- All Categories
- Featured Events
- Alumni
- Application Deadline
- Arts
- Campus Discourse
- Careers
- BU Central
- Center for the Humanities
- Charity & Volunteering
- Kilachand Center
- Commencement
- Conferences & Workshops
- Diversity & Inclusion
- Examinations
- Food & Beverage
- Global
- Health & Wellbeing
- Keyword Initiative
- Lectures
- LAW Community
- LGBTQIA+
- Meetings
- Orientation
- Other Events
- Religious Services & Activities
- Special Interest to Women
- Sports & Recreation
- Social Events
- Study Abroad
- Weeks of Welcome
- Day of ArafahAll day
- Eid-al-AdhaAll day
- IS&T RCS Tutorial - Introduction to Linux (Hands-on)10:00 am
- Virtual In-Studio AI Jam: Collaboratively Discovering AI11:00 am
- ECE PhD Prospectus Defense: Hee Jae Kim12:00 pm
- IS&T RCS Tutorial - Introduction to BU's Shared Computing Cluster (Hands-on)12:30 pm
- Tech Connect Drop In Hours1:00 pm
- Salsa On1, Footwork, Beg. 014:00 pm
- Hoop and Trapeze 015:00 pm
- Hoop and Trapeze 036:00 pm
- Aerial Silks Skills 077:00 pm
- Aerial Silks Skills 098:00 pm
- ECE Prospectus Defense: Mingyu Chen2:00 pm
- Aerial Silks Skills 114:00 pm
- Aerial Silks Skills 135:00 pm
- Aerial Silks Skills 155:15 pm
- Ballet, Beginner 016:00 pm
- Aerial Silks Intermediate 036:15 pm
- Aerial Silks Skills 177:30 pm
ECE Prospectus Defense: Mingyu Chen
ECE Prospectus Defense: Mingyu Chen
Title: Adaptive and Efficient RL Training: From Theoretical Foundations to LLM Post-Training
Presenter: Mingyu Chen
Advisor: Professor Xuezhou Zhang
Chair: Professor Ashok Cutkosky
Committee: Professor Xuezhou Zhang, Professor Wenchao Li, Professor Ashok Cutkosky, Professor Aldo Pacchiano
Google Scholar Link: https://scholar.google.com/citations?user=-C4-gdYAAAAJ&hl=en
Abstract: Reinforcement learning has become an important framework for both sequential decision-making and large language model post-training. However, many existing RL methods rely on problem-dependent prior knowledge, costly exploration, and expensive online optimization, limiting their scalability in complex environments and modern LLM applications. This work studies adaptive and efficient RL algorithms from both theoretical and practical perspectives. The first part develops parameter-free RL methods, which adapt to unknown quantities such as reward scales and state spaces, while maintaining strong regret guarantees without extensive manual tuning. The second part investigates efficient RL algorithms for LLM post-training, focusing on improving the sample and computational efficiency of RLHF and reasoning under sparse, skewed, or costly feedback. Together, this work aims to build principled RL methods that are both theoretically adaptive and practically scalable, bridging classical decision-making problems and modern language model alignment.
| When | 2:00 pm - 3:00 pm on 27 May 2026 |
|---|---|
| Building | PHO 339 |