Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation
作者: Henry Ascencio Trejo, Roel Pieters, Gokhan Alcan
分类: cs.RO, eess.SY
发布日期: 2026-09-09
备注: 20 pages, 19 figures
💡 一句话要点
提出自适应共享控制以解决人机协作中的非理性行为问题
🎯 匹配领域: 支柱一:机器人控制 (Robot Control)
关键词: 自适应控制 共享控制 有限理性 人机协作 动态规划 非线性系统 机器人技术
📋 核心要点
- 现有方法假设人类行为完全理性,未能有效应对实际中的有限理性行为,导致机器人辅助效果不佳。
- 论文提出了一种基于水平-k 有限理性模型的自适应共享控制方法,机器人根据观察到的人类行为动态调整其策略。
- 实验结果表明,所提方法在估计人类行为分布的准确性和机器人运行成本方面均优于传统的最大概率和概率加权策略基线。
📝 摘要(中文)
本研究考虑了非线性控制仿射系统中的自适应共享人机控制,放宽了完全理性人类的假设,机器人根据观察到的有限理性人类行为调整其辅助策略。我们使用两人博弈的水平-k 有限理性模型,通过交替最佳响应计算构建候选人类和机器人策略的有限库,并利用自适应动态规划近似相关的价值函数和策略。在共享控制交互中,状态转移残差比较了测量的系统演变与候选人类策略预测的轨迹。残差通过遗忘因子累积,并映射到有限候选库上的概率人类行为模型。机器人通过最小化预期合作成本计算分布感知的一步最佳响应,得到了一个闭式解。该方法在基准非线性系统稳定任务和二维操纵器共享控制设置中进行了评估,结果显示估计的人类行为分布与模拟分布之间的Kullback-Leibler散度降低,机器人在共享控制交互期间的累计运行成本低于基线策略。
🔬 方法详解
问题定义:本研究旨在解决在非线性控制仿射系统中,机器人如何有效适应有限理性人类行为的问题。现有方法通常假设人类行为完全理性,无法处理实际中人类的非理性决策,导致机器人辅助效果不理想。
核心思路:论文的核心思路是采用水平-k 有限理性模型,通过构建候选人类和机器人策略库,利用自适应动态规划来近似价值函数和策略,从而使机器人能够根据人类行为的实际表现动态调整其控制策略。
技术框架:整体架构包括几个主要模块:首先,通过交替最佳响应计算构建候选策略库;其次,利用状态转移残差比较实际系统演变与预测轨迹;最后,基于概率人类行为模型计算分布感知的一步最佳响应。
关键创新:最重要的技术创新在于引入了基于有限理性的人类行为模型,使机器人能够在共享控制中动态适应人类的非理性决策,而不是仅依赖于单一的候选策略或平均策略。
关键设计:在设计中,采用了遗忘因子来累积状态转移残差,并通过闭式解形式表达预期人类输入,确保了计算的高效性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,所提方法在估计人类行为分布时,Kullback-Leibler散度显著降低,且机器人在共享控制交互期间的累计运行成本低于传统基线策略,表明该方法在实际应用中具有更高的效率和适应性。
🎯 应用场景
该研究的潜在应用领域包括人机协作机器人、智能辅助设备和自动驾驶系统等。通过更好地理解和适应人类的非理性行为,机器人能够提供更有效的支持,提升人机交互的安全性和效率。未来,该方法可能在复杂环境中的自主决策和协作任务中发挥重要作用。
📄 摘要(原文)
This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.