Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
在多示例强化学习系统中整合公平性与可解释性
Abstract: Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions.
摘要: 利用教育交互数据预测学生表现,要求模型既要准确,又要具备足够的透明度以支持有效的干预措施;与此同时,人口统计学信息也引入了预测不公平的额外风险。
This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction.
本研究探讨了一个多目标框架,该框架结合了基于强化学习的多示例学习(RL-MIL)、对抗性去偏技术以及偏好条件超网络,用于预测处于风险中的学生。
MIL represents each student as a bag of weakly labeled interactions, while an RL agent selects informative instances for downstream classification.
多示例学习(MIL)将每位学生表示为一组弱标记的交互数据包,而强化学习(RL)智能体则负责为后续的分类任务选择具有信息量的示例。
Two hypernetwork variants are evaluated to determine whether a user-defined preference scalar can continuously control the trade-off between predictive performance and Equalized Odds.
研究评估了两种超网络变体,旨在确定用户定义的偏好标量是否能够持续控制预测性能与“均等几率”(Equalized Odds)之间的权衡。
The underlying RL-MIL baseline achieves strong classification performance, but both hypernetwork extensions exhibit mode collapse: changing the preference weight produces little systematic movement along the intended fairness-performance frontier.
底层的 RL-MIL 基准模型实现了强大的分类性能,但两种超网络扩展模型均出现了模式崩溃(mode collapse)现象:改变偏好权重几乎无法在预期的公平性-性能前沿面上产生系统性的移动。
The failure is associated with objective dominance, weak gradient propagation through the conditioning mechanism, and interactions between dynamically generated parameters.
这种失败与目标主导性、通过条件机制的梯度传播微弱,以及动态生成参数之间的相互作用有关。
The results show that fairness objectives can be incorporated into an interpretable RL-MIL pipeline, but preference conditioning alone does not guarantee controllable multi-objective behavior.
研究结果表明,公平性目标可以被整合到可解释的 RL-MIL 流水线中,但仅靠偏好条件设置并不能保证可控的多目标行为。
Robust fair RL-MIL therefore requires explicit mechanisms for gradient balancing, objective separation, and stability analysis.
因此,稳健的公平 RL-MIL 需要明确的梯度平衡、目标分离和稳定性分析机制。