ChatGPT and Gemini Rarely Agree on Top Local Businesses, Study Finds

ChatGPT and Gemini Rarely Agree on Top Local Businesses, Study Finds

研究发现:ChatGPT 和 Gemini 在推荐本地商家时鲜有共识

AI visibility is not a single score that a business can measure once and treat as settled. A cross-engine study of local-service searches found that ChatGPT and Gemini named the same top business in only 4.2% of identical queries. For small businesses trying to be discovered through AI assistants, that gap means a strong result in one engine may say very little about how another assistant presents the market. AI 的可见度并非一个可以一次性测量并定论的单一分数。一项针对本地服务搜索的跨引擎研究发现,在相同的查询中,ChatGPT 和 Gemini 推荐出同一家顶级商家的比例仅为 4.2%。对于试图通过 AI 助手被发现的小企业而言,这种差异意味着在一个引擎中获得良好的结果,并不能说明另一个助手会如何呈现该市场。

The research, published by Steady Demand in its AI Citation Ledger, examined 1,487 queries across 50 U.S. metropolitan areas and 10 service verticals. It focused on prompts such as “best plumber near me,” tracking the businesses named and the sources used to ground responses. Its central finding is practical: AI-driven discovery is fragmented by engine, source mix, location, and category. 这项由 Steady Demand 在其《AI 引用分类账》(AI Citation Ledger)中发布的研究,调查了美国 50 个大都市区和 10 个服务垂直领域的 1,487 次查询。研究重点关注“我附近的最佳水管工”等提示词,追踪了被提及的商家以及用于支撑回答的来源。其核心发现非常务实:AI 驱动的发现过程因引擎、来源组合、地理位置和类别的不同而呈现碎片化。

That does not prove that AI responses drive more leads than conventional local search. The study measures citations and top-name outcomes, not conversions or overall ranking quality. But it provides a useful baseline for marketers because it shows why checking a brand in one AI assistant is not enough to understand its broader AI visibility. 这并不能证明 AI 回答比传统的本地搜索能带来更多的潜在客户。该研究衡量的是引用情况和顶级推荐结果,而非转化率或整体排名质量。但它为营销人员提供了一个有用的基准,因为它表明,仅在一个 AI 助手中查看品牌表现,不足以了解其更广泛的 AI 可见度。

What the cross-engine data shows

跨引擎数据揭示了什么

The study compared how Gemini and ChatGPT answered the same local-business prompts. Their differences extended beyond the final recommendation. The systems often drew on different source ecosystems, which helps explain why they surface different businesses. 该研究对比了 Gemini 和 ChatGPT 如何回答相同的本地商家提示词。它们的差异不仅限于最终推荐结果。这两个系统通常依赖于不同的来源生态系统,这有助于解释为什么它们会呈现出不同的商家。

MeasureGeminiChatGPT
指标GeminiChatGPT
Exact top-business match between engines4.2% of identical queries produced the same top businessN/A
引擎间顶级商家完全匹配度4.2% 的相同查询产生了相同的顶级商家不适用
Typical citation mixAbout 60% of citations were business websitesMore reliance on Reddit and traditional directories
典型引用组合约 60% 的引用来自商家网站更依赖 Reddit 和传统目录网站
Overlap in cited domainsAbout 8% overlapN/A
引用域名的重叠度约 8% 的重叠不适用
Repeated-query source alignmentAbout 40% alignment, described as grounding driftN/A
重复查询的来源一致性约 40% 的一致性,被称为“基础漂移”(grounding drift)不适用
Top-result repeatability benchmarkAbout 7% top-match stability in AI-generated resultsNot specified separately in the supplied research
顶级结果的可重复性基准AI 生成结果中约 7% 的顶级匹配稳定性研究中未单独说明

The contrast with Google’s local pack is notable. In the study’s benchmark, Google local-pack top results reappeared about 90% of the time, compared with roughly 7% top-match stability for Gemini’s AI-generated results. That does not make one channel inherently better. It does show that an AI answer can be more variable than the familiar local-search result set businesses have historically monitored. 与谷歌本地搜索结果(Local Pack)的对比非常显著。在研究的基准测试中,谷歌本地搜索的顶级结果有约 90% 的时间会重复出现,而 Gemini 的 AI 生成结果的顶级匹配稳定性仅为 7% 左右。这并不意味着某个渠道天生更好,但它确实表明,AI 的回答比企业过去习惯监测的本地搜索结果集更具变数。

The source patterns also matter. Gemini’s heavier use of business websites suggests that clear, accessible on-site information can be especially important in its answers. ChatGPT’s greater use of Reddit and directories means third-party discussion and listing accuracy may have more influence there. Neither pattern supports a universal checklist, because the source mix changes by service category. 来源模式也很重要。Gemini 对商家网站的更多使用表明,清晰、易于访问的站内信息在其回答中尤为重要。而 ChatGPT 对 Reddit 和目录网站的更多使用意味着,第三方讨论和列表信息的准确性在那里可能具有更大的影响力。这两种模式都不支持通用的检查清单,因为来源组合会随服务类别而变化。

Why repeated prompts can produce different evidence

为什么重复的提示词会产生不同的证据

A second challenge is repeatability. When the researchers ran queries word for word again, cited sources aligned only about 40% of the time. The report calls this grounding drift: variation in the sources an AI system uses to support a response over repeated runs. For a business, this means a one-off screenshot of a favorable recommendation is weak evidence of durable visibility. The same applies to a negative or missing mention. 第二个挑战是可重复性。当研究人员逐字重复运行查询时,引用的来源仅有约 40% 的一致性。报告将其称为“基础漂移”(grounding drift):即 AI 系统在多次运行中用于支撑回答的来源存在差异。对于企业而言,这意味着一次性截取到的好评推荐,并不能作为持久可见度的有力证据。负面评价或缺失提及的情况也是如此。

A practical measurement approach for SMBs

针对中小企业的实用测量方法

The findings point toward a more disciplined way to evaluate AI discovery. Rather than asking whether a company “ranks in AI,” businesses can track whether they are consistently named and grounded across the assistants their customers may use. A practical monitoring process should include: 这些发现指向了一种更严谨的 AI 发现评估方式。企业不应再问“我在 AI 中排名如何”,而应追踪其是否在客户可能使用的各种助手中被持续提及并获得支撑。一个实用的监测流程应包括:

  • Test multiple engines separately: At minimum ChatGPT and Gemini when they are relevant to the audience. 分别测试多个引擎: 至少在与受众相关时,同时测试 ChatGPT 和 Gemini。
  • Use representative prompts: Reflect real customer language, services, locations, and decision stages. 使用具有代表性的提示词: 反映真实的客户语言、服务、地点和决策阶段。
  • Repeat the same prompts over time: To distinguish a one-off answer from a recurring pattern. 定期重复相同的提示词: 以区分一次性回答和重复出现的模式。
  • Record both the business named and the cited sources: Because sources reveal where each engine is finding evidence. 记录被提及的商家和引用的来源: 因为来源揭示了每个引擎从何处获取证据。
  • Segment results by service category and market: Since a tactic that helps one vertical or metro may not transfer to another. 按服务类别和市场细分结果: 因为对某个垂直领域或城市有效的策略,可能无法迁移到另一个领域。

This is not a call to chase every mention. It is a way to identify gaps that can be addressed with evidence-based work. If Gemini frequently cites business websites but a company has thin service pages or inconsistent location details, improving those pages may make its information easier to ground. If ChatGPT repeatedly draws on directories or community discussions, the business can review whether core listings are accurate and whether public information about its services is clear. 这并不是呼吁去追逐每一次提及,而是一种识别差距的方法,以便通过基于证据的工作来解决问题。如果 Gemini 频繁引用商家网站,但某公司的服务页面内容单薄或地点信息不一致,那么改进这些页面可能会使其信息更容易被采纳。如果 ChatGPT 反复引用目录网站或社区讨论,企业则可以检查其核心列表信息是否准确,以及关于其服务的公开信息是否清晰。

Structured data can be part of that work when it accurately represents the business and its services. However, the study does not establish that a particular markup implementation, directory, or content tactic guarantees inclusion in either assistant. Its evidence supports monitoring the source environment, not promising a universal optimization formula. 当结构化数据能准确代表企业及其服务时,它可以成为这项工作的一部分。然而,该研究并未证实某种特定的标记实现、目录或内容策略能保证被任何一个助手收录。其证据支持的是监测来源环境,而不是承诺某种通用的优化公式。

For marketing and customer-support teams, the same principle applies to research workflows. AI assistants can be useful for discovering how consumers may encounter a brand, but their responses should not be treated as a stable market ranking or as a substitute for verified local-search data. Teams should document the engine, prompt, location, date, cited sources, and repeat runs when using outputs in decisions. 对于营销和客户支持团队来说,同样的原则也适用于研究工作流程。AI 助手有助于发现消费者如何接触到品牌,但其回答不应被视为稳定的市场排名,也不应替代经过验证的本地搜索数据。团队在将这些输出用于决策时,应记录引擎、提示词、地点、日期、引用来源以及重复运行的情况。