Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

迈向阿片类药物使用障碍治疗中,预测治疗留存率与过早中断的机器学习模型的公平性

Abstract: Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support.

摘要: 阿片类药物使用障碍(MOUD)药物治疗中持续存在的低留存率和低完成率,促使人们利用机器学习(ML)模型来预测留存情况,并识别有治疗过早中断风险的患者。然而,这些模型在不同患者群体间的公平性在很大程度上仍未得到探索,这引发了人们对其在治疗决策支持中应用的担忧。

This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set-Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD.

本研究系统地评估了用于预测 MOUD 留存率和过早中断的机器学习模型的算法公平性,并调查了偏差缓解技术的有效性。我们使用了横截面治疗事件数据集(TEDS-D),该数据集包含了 2015 年至 2019 年间美国出院患者的治疗记录,并训练了四个机器学习模型,旨在预测接受门诊 MOUD 治疗的个体过早中断治疗的情况以及超过 180 天的留存情况。

We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed bias mitigation techniques and their effects on both fairness and predictive performance.

我们评估了按种族、族裔、年龄和性别定义的患者子群体的整体性能和子群体层面的错误率,并辅以模型解释分析。我们进一步评估了偏差缓解技术及其对公平性和预测性能的影响。

Our findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.

研究结果表明,即使在整体预测性能看似可接受的情况下,用于 MOUD 结果预测的机器学习模型仍可能表现出子群体层面的性能差距;此外,偏差缓解技术可以在不产生权衡的情况下减少,但无法完全消除这些差距。通过证明公平性感知评估和透明报告子群体性能的重要性,本研究为在 MOUD 治疗环境中负责任且结合具体情境地使用机器学习模型进行风险分层和护理优先级排序提供了实践见解。