Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
完成任务还不够:评估生成式智能体在累积挑战下的韧性与协作参与度
Abstract: Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time.
摘要: 生成式 AI 智能体的持续部署不仅要求其能够完成孤立的任务。智能体必须在重复的交互、不断变化的条件以及共享工作流中对人的依赖关系中保持有效性,尤其是在技术、人为和操作层面的干扰随时间累积的情况下。
We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge.
我们提出了“操作韧性”(operational resilience)和“协作参与度”(considerate participation)作为评估此类智能体的两个互补维度:前者衡量智能体如何在受阻时恢复工作,同时保持进度并沟通其能力边界;后者则衡量智能体在适应过程中如何顾及受影响的人员、角色边界以及周边工作流。然而,在累积挑战的背景下,这两个方面仍未得到充分研究。
We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates.
我们针对两种生成式 AI 模型和十二项源自利益相关者的任务,模拟了 120 条医疗保健工作轨迹,并设置了轻度、中度和重度挑战。我们对比了文本行动计划、提示词引导的内部评估,以及定量的结构化工作负载与情感报告,以考察随着挑战的累积,智能体的行为及其报告的状态如何发生变化。
Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments.
在操作韧性方面,智能体从自主恢复转向对人类的更多依赖;虽然在结构化报告中反映出工作负载增加和负面情绪,但在文本回复中却很少表达压力。在协作参与度方面,智能体的适应方式从单纯的任务导向扩展到任务重构、关注他人、角色边界调整以及更广泛的协调,并在行动和内部评估中表现出明显的模式差异。
From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.
基于这些发现,我们总结了五个部署困境,涉及持久性、注意力、角色边界、状态披露和升级机制,这些问题需要利益相关者明确界定,并为学习、情境化评估和具身适应等方面的技术实现提供了进一步的参考。