HEIR: Learning Human-Entity Interactions with Functional Roles

HEIR: Learning Human-Entity Interactions with Functional Roles

HEIR:学习具有功能角色的“人-实体”交互

Understanding human-entity interactions requires recovering each person-action event’s participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. 理解“人-实体”交互(Human-Entity Interactions)需要恢复每个人物-动作事件中的参与者、角色以及共享身份。这种结构能够通过明确“谁对哪些实体采取了什么行动”来支持具身智能体,从而为共享环境中的预测和协作提供信息。

Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. 标准的 HOI(人-物交互)指标通常只对单个链接进行评分,导致对完整事件构成的评估不足。我们引入了 HEIR(具有功能角色的“人-实体”交互),这是一个针对物体交互、人际交互和自我导向交互中完整且已标注的参与者-角色集合的图像基准。

It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. 该基准包含 18,730 张图像、6 种角色、105 种动作和 437 个名词,涵盖了共享实体、角色转换和重复填充项;其中 51.6% 的图像包含多个参与者,62.1% 的图像包含多个动作。HEIR 将关系 AP(平均精度)与完整集合 AP 及结构化评估相结合。

We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. 我们还引入了 CoRISP(组合式角色感知交互集合预测),它利用共享的实体身份来整合角色条件下的证据,并预测归一化的参与者-角色集合。基数和角色多重性势能通过事件规模和角色构成来耦合分配,并实现了每个事件的精确归一化。

Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. 在 16 个基准模型中,即使在对齐动作权重后,关系排名与完整事件排名仍存在差异。CoRISP 在 HEIR 的重复角色事件和共享参与者图像任务中表现领先,Set mAP 分别提升了 2.87 和 3.82 个百分点。

On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at this https URL. 在 V-COCO 数据集上,CoRISP 在场景 1/2 下的双槽动作中分别达到了 73.72/76.23 的角色 AP 和 61.06/68.59 的完整集合 AP。这些结果证明了在评估单个关系的同时,学习和评估事件构成具有重要价值。代码和数据集已在链接中公开。