Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

基于即时成本条件风险价值(CVaR)且具备性能保证的在线 POMDP 风险规避规划

Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk within the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners.

摘要: 在线部分可观测马尔可夫决策过程(POMDP)规划器通常优化期望累积成本,但当置信度(belief)在某些高成本状态上分配了较大权重时,这种方法可能会掩盖危险状态。现有的风险规避方法将静态或动态条件风险价值(CVaR)应用于价值函数,以捕捉轨迹层面的风险,但这些方法存在两个缺陷:(i) 由于将即时成本保留为置信度下状态相关成本的期望值,置信度内部的风险未得到解决;(ii) 由于修改了价值函数,它们需要专门定制的新算法,而无法直接复用现有的基于期望的规划器。

We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation.

我们转而将 CVaR 应用于每一步置信度下的即时成本,直接针对当前状态的单步不确定性。由于保留了标准的期望累积回报作为目标,所得问题具有标准的 MDP 结构:只需更改成本计算方式,任何基于期望的 POMDP 规划器都可以具备风险敏感性。

We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.

我们继承了策略评估和稀疏采样的有限时间保证(且估计误差与风险水平无关),并作为核心理论成果,证明了粒子置信度 MDP 代理与原始 POMDP 之间差距的有限时间界限。这些共同构成了从真实 POMDP 值到算法估计值的端到端保证。在风险中性极限下,该公式可还原为标准的基于期望的规划方法。