What Are RAG Agents? A Plain-English Guide for Business Owners

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

Originally published at jy-labs.com. Updated 2026-10-07. RAG agents (Retrieval-Augmented Generation) are AI assistants that search your company’s own documents before they answer, then cite the source, so the answer comes from your policy or contract instead of the model’s general training. This guide covers how they work in October 2026, when a per-seat tool like Copilot or Claude Team is enough, when a custom build pays off, and what each option costs.

该文章最初发布于 jy-labs.com,更新于 2026 年 10 月 7 日。RAG 智能体(检索增强生成)是一种人工智能助手,它在回答问题前会先搜索贵公司的内部文档,并引用来源,从而确保答案基于您的政策或合同,而非模型的一般性训练数据。本指南涵盖了它们在 2026 年 10 月的工作原理、何时使用 Copilot 或 Claude Team 等按席位付费的工具已足够、何时定制开发更划算,以及每种选项的成本。

The problem Every assistant now claims to know your business. Microsoft 365 Copilot, ChatGPT Business, Claude Team, and Google Workspace each ship connectors to your files, and every AI vendor calls their product an agent. An owner comparing a $20 per seat subscription with a five-figure custom quote cannot tell what the extra money buys, or whether a 1 million token context window makes retrieval unnecessary. This guide explains the mechanics in plain language, lists verified October 2026 prices, and gives you a test for which option fits your document library.

问题所在:现在每个助手都声称了解您的业务。Microsoft 365 Copilot、ChatGPT Business、Claude Team 和 Google Workspace 都提供了连接到您文件的接口,且每家 AI 供应商都将其产品称为“智能体”。当企业主在比较每月 20 美元的按席位订阅和五位数的定制报价时,无法分辨多出的费用买到了什么,也无法判断 100 万 token 的上下文窗口是否让检索变得多余。本指南用通俗易懂的语言解释了其机制,列出了 2026 年 10 月经核实的定价,并为您提供了一个测试,以判断哪种选项适合您的文档库。

The approach What RAG means RAG stands for Retrieval-Augmented Generation. A language model on its own answers from what it learned in training. A RAG agent runs a search over your documents first, hands the matching passages to the model, and the model writes an answer from those passages with a citation back to the page it used. Three parts do the work: Ingestion. Someone collects your PDFs, Word files, wiki pages, tickets, and emails, splits them into passages, and indexes them. In most builds, indexing means creating embeddings (numeric fingerprints of meaning) and storing them in a vector database, plus a keyword index for exact terms like part numbers and case citations. Retrieval. When an employee asks a question, the system searches the index and pulls the handful of passages most likely to contain the answer. Generation. The model reads those passages and writes the answer, quoting the source. The word “agent” means the model decides what to search for, runs more than one search when the first comes back thin, and calls other tools (your CRM, your ticketing system) when the answer lives there. A plain RAG pipeline searches once. An agent keeps searching until it has what it needs or reports it could not find the answer.

方法:RAG 的含义。RAG 代表检索增强生成。语言模型本身是根据其训练中学到的知识来回答问题的。而 RAG 智能体会先对您的文档进行搜索,将匹配的段落交给模型,模型再根据这些段落撰写答案,并注明所引用的页面。其工作分为三个部分:摄入。将您的 PDF、Word 文档、维基页面、工单和电子邮件收集起来,拆分成段落并建立索引。在大多数构建中,索引意味着创建嵌入(含义的数字指纹)并将其存储在向量数据库中,同时为零件编号和案例引用等精确术语建立关键词索引。检索。当员工提出问题时,系统会搜索索引并提取最可能包含答案的几段内容。生成。模型阅读这些段落并撰写答案,同时引用来源。“智能体”一词意味着模型可以决定搜索什么,当第一次搜索结果不足时会进行多次搜索,并在答案存在于其他工具(如您的 CRM 或工单系统)中时调用这些工具。普通的 RAG 流水线只搜索一次,而智能体会持续搜索,直到获得所需信息或报告无法找到答案为止。

How a RAG agent differs from ChatGPT or Claude out of the box In 2025 the answer was simple: ChatGPT knew the internet and a RAG agent knew your files. In October 2026 that line has moved. Claude ships connectors to Google Drive, Microsoft 365, Slack, Notion, and HubSpot on every plan, and enterprise search across company data on Team and Enterprise. ChatGPT Business supports connectors and custom MCP servers for company knowledge. Microsoft 365 Copilot reads your SharePoint and OneDrive by default. So the question for an owner becomes who controls the four things a RAG agent gets right or wrong: Permissions. Does the system respect who is allowed to see which document, or does a junior hire get answers drawn from the partner-only folder? Retrieval quality. Does it find the right clause in a 90-page contract, or the first paragraph with a matching keyword? Citations. Does every answer link to the page it came from, so a human verifies in one click? Freshness. When you update the return policy, does the agent answer from the new version the same day? A per-seat assistant gives you reasonable defaults on all four and no control over any of them. A custom build gives you control and the bill that comes with it. The next section prices both.

RAG 智能体与开箱即用的 ChatGPT 或 Claude 有何不同?2025 年时答案很简单:ChatGPT 了解互联网,而 RAG 智能体了解您的文件。到了 2026 年 10 月,这条界限已经模糊。Claude 在所有计划中都提供了连接 Google Drive、Microsoft 365、Slack、Notion 和 HubSpot 的接口,并在 Team 和 Enterprise 计划中提供跨公司数据的企业搜索。ChatGPT Business 支持连接器和用于公司知识的自定义 MCP 服务器。Microsoft 365 Copilot 默认读取您的 SharePoint 和 OneDrive。因此,对于企业主来说,问题变成了谁来控制 RAG 智能体在以下四个方面表现的好坏:权限。系统是否尊重谁有权查看哪些文档,还是说初级员工也能从仅限合伙人查看的文件夹中获取答案?检索质量。它是在 90 页的合同中找到正确的条款,还是仅仅找到第一个包含匹配关键词的段落?引用。每个答案是否都链接到其来源页面,以便人工一键核实?时效性。当您更新退货政策时,智能体是否能在当天根据新版本回答?按席位付费的助手在上述四点上提供了合理的默认设置,但您无法控制任何一项。定制开发则赋予您控制权,但也伴随着相应的费用。下一节将对两者进行定价分析。

Three ways to get RAG in October 2026, with prices Verified list prices as of this week. They move, so check before you budget. 1. A per-seat assistant with connectors. The cheapest way to find out whether your team will use this. Microsoft 365 Copilot: $30 per user per month paid yearly. The small-business add-on, Microsoft 365 Copilot Business, lists at $21 per user per month on an annual term, discounted to $18 through December 31, 2026 for the first year, capped at 300 users, and requires a qualifying Microsoft 365 plan. Claude Team: $20 per seat per month billed annually, $25 monthly, 2 to 150 seats. Connectors on every plan, enterprise search on Team and Enterprise. ChatGPT Business: $20 per user per month billed annually, $25 monthly, two-seat minimum, with connectors and support for custom MCP servers. Google Workspace: Standard at $14 per user per month includes Gemini in Gmail, Docs, Sheets, and Drive plus Gemini Notebook. For a 20-person office, that is $280 to $600 per month. If your documents already live in Microsoft 365 or Google Drive and your questions are “what does the handbook say about PTO,” start here.

2026 年 10 月获取 RAG 的三种方式及价格(本周核实的价格,价格会有变动,预算前请核实):1. 带连接器的按席位助手。这是了解您的团队是否会使用此功能的最便宜方式。Microsoft 365 Copilot:每用户每月 30 美元,按年支付。小型企业附加组件 Microsoft 365 Copilot Business 年度条款定价为每用户每月 21 美元,2026 年 12 月 31 日前首年优惠至 18 美元,上限 300 用户,且需要符合条件的 Microsoft 365 计划。Claude Team:每席位每月 20 美元(按年计费)或 25 美元(按月计费),2 至 150 个席位。所有计划均含连接器,Team 和 Enterprise 计划含企业搜索。ChatGPT Business:每用户每月 20 美元(按年计费)或 25 美元(按月计费),至少两个席位,支持连接器和自定义 MCP 服务器。Google Workspace:标准版每用户每月 14 美元,包含 Gmail、Docs、Sheets 和 Drive 中的 Gemini 以及 Gemini Notebook。对于 20 人的办公室,每月费用为 280 至 600 美元。如果您的文档已存储在 Microsoft 365 或 Google Drive 中,且您的问题是“手册中关于带薪休假(PTO)是怎么说的”,请从这里开始。

  1. A managed retrieval service inside a custom app. You get your own interface, your own rules, and the vendor runs the index. OpenAI file search: $0.10 per GB of index per day with the first GB free, plus $2.50 per 1,000 searches. Pinecone vector database: free Starter tier up to 2 GB, then a $50 per month minimum on Standard, with storage at $0.33 per GB per month. pgvector: open-source vector search inside Postgres 13 and later. On Supabase that runs on the free tier or the $25 per month Pro plan. For a small document library this is the cheapest index you will find. Model tokens: Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output. Claude Haiku 5.5 costs $0.10 and $0.50 for prompts under 100,000 tokens. A question that retrieves 20,000 tokens of passages and writes a 500-token answer costs about 4.5 cents on Sonnet and under a quarter of a cent on Haiku. One thousand questions a month is $45 or about $2.25. Running costs for a small business land under $100 per month in most cases. The money is in the build.

  2. 自定义应用内的托管检索服务。您拥有自己的界面和规则,由供应商运行索引。OpenAI 文件搜索:每天每 GB 索引 0.10 美元(首 GB 免费),外加每 1,000 次搜索 2.50 美元。Pinecone 向量数据库:Starter 层级 2GB 以内免费,之后 Standard 层级每月最低 50 美元,存储费为每月每 GB 0.33 美元。pgvector:Postgres 13 及更高版本中的开源向量搜索。在 Supabase 上,它可以在免费层级或每月 25 美元的 Pro 计划中运行。对于小型文档库,这是您能找到的最便宜的索引。模型 Token:Claude Sonnet 5.5 每百万输入 Token 2 美元,每百万输出 Token 10 美元。Claude Haiku 5.5 在提示词低于 10 万 Token 时分别为 0.10 美元和 0.50 美元。一个检索 2 万 Token 段落并撰写 500 Token 答案的问题,在 Sonnet 上成本约为 4.5 美分,在 Haiku 上不到 0.25 美分。每月 1,000 个问题费用为 45 美元或约 2.25 美元。小型企业的运行成本在大多数情况下每月不到 100 美元。主要的投入在于开发构建。

  3. A custom build. Someone designs the ingestion, picks the embedding model, writes the permission layer, builds the evaluation set, and connects your systems. JY Labs scopes a typical deployment at 6 to 7 weeks, with a working prototype by week 3 and 8 to 10 weeks wh

  4. 定制开发。由专人设计摄入流程、选择嵌入模型、编写权限层、构建评估集并连接您的系统。JY Labs 评估一个典型的部署周期为 6 到 7 周,第 3 周可提供工作原型,8 到 10 周……