Abstract
We investigate a novel safe reinforcement learning problem with step-wise violation constraints. Our problem differs from existing works in that we focus on stricter step-wise violation constraints and do not assume the existence of safe actions, making our formulation more suitable for safety-critical applications that need to ensure safety in all decision steps but may not always possess safe actions, e.g., robot control and autonomous driving. We propose an efficient algorithm SUCBVI, which guarantees an (O) over tilde(root ST) or gap-dependent (O) over tilde (S/C-gap + S(2)AH(2)) step-wise violation and an (O) over tilde(root H(3)SAT) regret. Lower bounds are provided to validate the optimality in both violation and regret performance with respect to the number of states S and the total number of steps T. Moreover, we further study an innovative safe reward-free exploration problem with step-wise violation constraints. For this problem, we design the algorithm SRF-UCRL to find a near-optimal safe policy, which achieves a nearly state-of-the-art sample complexity (O) over tilde((S(2)AH(2)/epsilon + H(4)SA/epsilon(2))(log(1/delta) + S)), and guarantees an (O) over tilde(root ST) violation during exploration. Experimental results demonstrate the superiority of our algorithms in safety performance and corroborate our theoretical results.