Characterizing Human-Likeness in AI Generated Poetry: A Zero-shot Classification Study
Characterizing Human-Likeness in AI Generated Poetry: A Zero-shot Classification Study
刻画人工智能生成诗歌的类人特征:一项零样本分类研究
Abstract: With the advancement of AI technologies, Generative AI (GenAI) and human written text have become nearly indistinguishable. Additionally, the global standardization of AI chatbots made academic malpractice more frequent. Furthermore, existing research indicates GenAI poems are the most difficult to distinguish even without any modification thus, GenAI poems are naturally deemed human-like by modern detectors.
摘要: 随着人工智能技术的进步,生成式人工智能(GenAI)生成的文本与人类创作的文本已几乎无法区分。此外,人工智能聊天机器人的全球标准化使得学术不端行为愈发频繁。现有研究表明,即使不做任何修改,GenAI 创作的诗歌也最难被辨别,因此现代检测器往往自然地将 GenAI 诗歌视为具有“类人”特征。
However, the objectivity of such dissertations needs to be verified against modern detection tools but the subjectivity of poetry and the black-box nature of the modern LLMs (Large Language Models) architectures made verification of such work quite complicated. Hence, the main objective of the research is to deduce the attributes of English poetry that contribute classification and misclassification of both human and AI poems and provide corroborating or contradicting evidence to the poetry distinguishability claim.
然而,此类论点的客观性需要通过现代检测工具进行验证,但诗歌的主观性以及现代大语言模型(LLM)架构的“黑箱”特性,使得相关验证工作变得相当复杂。因此,本研究的主要目标是推导出导致人类诗歌与 AI 诗歌被正确分类或误分类的英语诗歌属性,并为“诗歌可辨别性”这一主张提供支持或反驳的证据。
For such characterizations, we propose a Zero-shot detection pipeline with a dataset consisting of both human and AI poems to verify the distinguishability of human and AI creation and extract the aforementioned crucial attributes for accurate classification. Extraction of such attributes provides benefits in two ways: firstly, it reduces the margin of training needed as only the poems based on misclassifying attributes need to be trained and fine tuned and finally provides a critical insight to the GenAI detection dilemma to strengthen the modern detection pipelines.
为了实现这些特征刻画,我们提出了一个零样本(Zero-shot)检测流程,并使用包含人类和 AI 诗歌的数据集来验证两者创作的可辨别性,同时提取上述用于准确分类的关键属性。提取这些属性具有双重益处:首先,它减少了所需的训练量,因为只需针对导致误分类的属性相关的诗歌进行训练和微调;其次,它为解决 GenAI 检测难题提供了关键见解,从而加强了现代检测流程。