What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews
What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews
用户如何看待生成式 AI:基于应用商店评论的跨平台 NLP 信任与摩擦分析
Abstract: Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store.
摘要: 生成式 AI (GenAI) 应用已实现快速的消费者普及,但目前鲜有大规模研究探讨用户感知的质量、信任度及采用障碍。我们对六款主流 GenAI 应用(ChatGPT、Gemini、Microsoft Copilot、Claude、DeepSeek 和 Perplexity)的应用商店评论进行了首次跨应用分析之一,研究涵盖了来自 Google Play 和 Apple App Store 的 17,012 条英文评论。
We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews.
我们结合了 BERTopic 主题建模与 RoBERTa 情感分类,并使用卡方检验、Kruskal-Wallis 检验以及带有 Bonferroni 校正的多项逻辑回归来评估跨应用之间的差异。上述两个分析组件均通过 300 条评论的分层样本进行了人工编码验证。
Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications.
结果显示,负面情绪主要集中在广告(91%)、身份验证(89%)、服务器稳定性(83%)和订阅定价(73%)方面。不同应用之间的情感差异显著,其中 Claude 的负面情绪比例最高(47.7%),但同时拥有一个非常热情的用户群体,这表明存在统计学意义上的两极分化。尽管各应用的评论数量不等,但这些发现依然具有稳健性。
As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.
作为探索性观察,部分 DeepSeek 的评论提出了与其中国背景相关的地缘政治和数据隐私担忧;同时,我们提出的“信任摩擦评分”(Trust Friction Score)将各应用特有的信任与可用性障碍总结为可解释的维度。本研究为消费者生成式 AI 应用中的用户信任、可用性及采用障碍提供了经过验证且具可操作性的证据。