| |
Persistent Safety Set Guided Offline Safe Reinforcement Learning Speaker:  Brahmanage Janaka Chathuranga THILAKARATHNA Ph.D. Candidate School of Computing and Information Systems Singapore Management University
| Date: Time: Venue: | | 11 August 2026, Tuesday 11:30am – 12:00pm Meeting room 4.4, Level 4. School of Computing and Information Systems 1, Singapore Management University, 80 Stamford Road Singapore 178902
Please register by 10 August 2026. 
|
|
About the Talk Offline safe reinforcement learning learns high-return policies that satisfy hard safety constraints using only a pre-collected dataset. This setting is challenging due to the inability to explore, and the risk of propagating value errors through unsafe state-space regions. To address this, first, we characterize the safe state region by developing a framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator, yielding a persistent safety set, from which the agent can remain safe indefinitely. Second, we show that several existing safety set estimation methods (e.g., reachability-constrained RL) can be formulated within our CBF learning framework, highlighting its generality. We further propose a new CBF that ensures safety under environment dynamics uncertainty, unlike standard CBFs designed for deterministic settings. Third, we propose a new reward maximization algorithm that effectively exploits our learned persistent safety set for reward critic estimation. Empirical results on standard benchmarks show that our approach achieves state-of-the-art safety with fewer constraint violations while maintaining competitive returns.
This is a Pre-Conference talk for The 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026). About the Speaker Janaka Brahmanage is a fourth-year PhD candidate in Computer Science, conducting research under the guidance of Associate Prof. Akshat Kumar at the SMU School of Computing and Information Systems. His research focuses on safe reinforcement learning and multi-agent systems.
|