I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

I-CARE:文本到图像模型在可控、多样且具有代表性的机器遗忘设置中对干扰相关现象的分析

Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been retained (henceforth, interference) remains poorly characterized and inconsistently evaluated.

摘要: 机器遗忘(Machine unlearning)旨在研究如何从人工智能模型中移除知识,使系统“忘记”其先前学习过的概念。尽管生成式机器遗忘领域进展迅速,但对于那些本应保留的语义相关概念所遭受的意外退化(以下简称“干扰”),目前仍缺乏充分的表征和一致的评估。

This paper introduces I-CARE, a methodology that formalizes interference as a first-class object of study in generative unlearning. Rather than proposing a new benchmark or unlearning algorithm, I-CARE provides formal definitions for tasks, metrics, and templates for reporting results, enabling the systematic and reproducible study of interference across unlearning settings.

本文介绍了 I-CARE,这是一种将“干扰”正式化为生成式遗忘研究中核心对象的分析方法。I-CARE 并非旨在提出新的基准测试或遗忘算法,而是为任务、指标和结果报告模板提供了正式定义,从而能够在不同的遗忘设置中对干扰进行系统且可复现的研究。

While our methodology is designed to remain valid as models and unlearning algorithms evolve, decoupling long-term scientific insight from transient empirical results, we present a feasibility demonstration with state-of-the-art algorithms and frequently used datasets. The results demonstrate that I-CARE enables meaningful analysis of interference patterns across multiple unlearning settings, establishing the practical applicability of the framework.

虽然我们的方法旨在随着模型和遗忘算法的演进而保持有效性,从而将长期的科学洞察与短暂的实证结果解耦,但我们仍通过最先进的算法和常用数据集展示了其可行性。结果表明,I-CARE 能够对多种遗忘设置中的干扰模式进行有意义的分析,确立了该框架的实际应用价值。

The software implementation of the methodology is provided in an open-source framework, together with a web-based graphical interface that enables exploration of the outcomes of this study without requiring direct interaction with the codebase or specialized data analysis tools.

该方法的软件实现已通过开源框架提供,并配备了基于 Web 的图形界面。用户无需直接与代码库交互或使用专业的数据分析工具,即可探索本研究的成果。