What Would A Serious AI Product Look Like?

What Would A Serious AI Product Look Like?

严肃的 AI 产品应该是什么样的?

Every “AI” tool is missing critical features that you would need if you wanted to do real work with them. What would those look like, if they existed? 目前的每一个“AI”工具都缺少你进行实际工作时所必需的关键功能。如果这些功能存在,它们会是什么样子?

One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like. 我对当前这一代“AI”产品的一个不满之处在于,它们似乎并没有认真对待自己的前提。我看着大量卑躬屈膝的聊天机器人声称自己是严肃的问题解决工具,我心想,这根本不是解决问题的工具该有的样子。

Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition as a product that makes me feel, constantly, whenever I am interacting with them, that they are less a software product than that they are a grift, a scam designed to make me feel like I am interacting with a product that has capabilities that it simply does not, to try to lull me into a false sense of security that I can trust it. 在谈论前沿实验室巨大的伦理问题之前,仅从它们作为产品的构成来看,就让我感到——每当我与它们交互时——它们与其说是软件产品,不如说是一场骗局。这种骗局旨在让我觉得我正在使用一个具备某种能力的产品,而实际上它根本不具备,从而试图让我产生一种可以信任它的虚假安全感。

The frontier labs are of course the worst offenders, but every criticism here applies just as much to Ollama, which (if anything, due to the obviously poorer quality of the available models themselves) needs these features even more than the frontier labs do. Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it. 前沿实验室当然是罪魁祸首,但这里的每一条批评同样适用于 Ollama。甚至可以说,由于其可用模型本身的质量明显较差,Ollama 比前沿实验室更需要这些功能。在此,我将列出一些功能,如果具备这些功能,或许能让我相信一个基于大语言模型(LLM)的产品——特别是专注于研究或软件开发的产品——是真心实意地在帮助我完成有用的工作。

Make “Checking For Mistakes” A First-Class Feature

将“检查错误”作为核心功能

This is the biggest issue, and the major reason that I was inspired to write this post. It is a truth universally acknowledged, that AIs cannot reliably provide information. I could cite a ton of news articles and studies about this fact, but there is no need. Every single chatbot admits this, up front, in a fine-print disclaimer as a core part of their user interface. 这是最大的问题,也是我写这篇文章的主要原因。众所周知,AI 无法可靠地提供信息。我可以引用大量新闻报道和研究来证明这一点,但没必要。每一个聊天机器人都开诚布公地在用户界面的核心位置,以细小的免责声明承认了这一点。

Gemini says “AI can make mistakes, so double-check responses”, Claude says “Claude is AI and can make mistakes. Please double-check responses.” ChatGPT says “ChatGPT can make mistakes. Check important info.”. Every time I see that last one, I wonder how I’m supposed to know what “info” is supposed to be “important”. All of these warnings are all small, gray text, painfully obviously included as legalese to push responsibility back onto the user rather than to help with anything. Gemini 说“AI 可能会出错,请仔细核对回复”;Claude 说“Claude 是 AI,可能会出错。请仔细核对回复”;ChatGPT 说“ChatGPT 可能会出错。请检查重要信息”。每次看到最后那句,我都在想,我怎么知道哪些信息才算“重要”?所有这些警告都是灰色的细小文字,显然是为了推卸责任而加入的法律条文,而不是为了提供任何实际帮助。

This is a core limitation of all these products. Checking their output is a part of the workflow for using them that: you absolutely cannot skip or skimp on without creating risks to yourself and whoever you are conveying its output to, and, it is very easy to skip or skimp on and you are encouraged at every turn to do so, because “just trust the output” is one of the quickest ways to save time. 这是所有这些产品的核心局限性。检查它们的输出是使用流程的一部分,你绝对不能跳过或敷衍,否则会给自己以及任何接收这些输出的人带来风险。然而,跳过或敷衍又非常容易,而且系统在每一步都在诱导你这样做,因为“直接信任输出”是节省时间最快的方法之一。

A chatbot product that took this weakness seriously, as an actual consideration for using it, would put a checkbox next to every claim in its output. It would be a 2-column worksheet, where you’ve got the LLM output in the first column, and next to it, human notes in the second column, explaining what work went into checking this claim, and a big checkbox that you would only check off after you believe you’d checked its claims thoroughly enough. 如果一个聊天机器人产品认真对待这一弱点,并将其作为使用时的实际考量,它会在输出的每一个论点旁边放置一个复选框。它应该是一个双栏工作表:第一栏是 LLM 的输出,第二栏是人类的备注,解释为核实该论点所做的工作,以及一个只有在你认为已经彻底核实了其论点后才会勾选的大复选框。

Coding assistants would need to have some version of this as well. Right now, this is pushed off into code review, which means it is a dark pattern which subtly encourages the “author” to offload this work to their code reviewer without ever looking. Once again, “it’s probably fine, I don’t need to check” is the quickest way to save time and churn out those PRs faster. 编程助手也需要具备类似的功能。目前,这项工作被推给了代码审查,这意味着这是一种暗黑模式,它微妙地鼓励“作者”在不看代码的情况下将这项工作转嫁给代码审查员。再一次,那句“应该没问题,我不用检查”是节省时间并更快产出 PR 的最快方式。

It might even be useful for coding harnesses to have some affordance for checking code before it even runs tests. As the vendors themselves have admitted, it’s not just expensive to burn tokens on your “AI”, you also end up burning far more compute on the AI. Being able to check your diffs before sending them over to uselessly exhaust your testing compute cluster would be useful. If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously. 如果编程工具能在运行测试之前就提供检查代码的手段,那将会非常有用。正如供应商自己所承认的,在“AI”上消耗 Token 不仅昂贵,而且最终会在 AI 上消耗更多的计算资源。在将代码差异(diffs)发送出去并无谓地耗尽测试计算集群之前,能够先进行检查是非常有用的。如果你的产品告诉我它会出错,并且必须由我来检查错误,但却不给我提供任何检查错误的工具,那我就无法认真对待它。

More Citations to Check, And More Details

提供更多可供核实的引用和细节

Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to “get bored” halfway through the list and simply stop including citations at some point. When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it’s barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side. This is backwards. 大多数聊天机器人倾向于直接给出答案,而不是引用来源。在我个人的使用中,我发现当要求它们提供带有明确来源的引用列表时,它们似乎会在列表写到一半时“感到厌倦”,并在某个点直接停止提供引用。当机器人确实包含引用时,它们通常以行内注释的形式呈现,只显示搜索结果的域名,字体小到几乎无法辨认,图标也小到不足 16 像素,同样难以辨认。这完全是本末倒置。

Now, I am aware that these citations do come from somewhere, and in an attempt to reduce hallucinations, all of the major providers support some form of “grounding”, and that those little barely-readable citation links are referencing actual structures in the RAG pipeline and not just potentially-hallucinated tokens, but I’m not talking about the underlying machinery in the model, I’m talking about the presentation to the user. 我当然知道这些引用是有出处的,为了减少幻觉,所有主要供应商都支持某种形式的“溯源”(grounding),那些难以辨认的小引用链接确实指向了 RAG(检索增强生成)流水线中的实际结构,而不仅仅是潜在的幻觉 Token。但我谈论的不是模型底层的机制,而是呈现给用户的方式。

Plus, regardless of whether a snippet of text came from a RAG query, we know that LLMs can never provide an authoritative result; it’s a fundamental limitation of the technology. They can still garble the results of RAG as much as they can misrepresent any other training data. This means that it must never present its results as authoritative. 此外,无论一段文本是否来自 RAG 查询,我们都知道 LLM 永远无法提供权威的结果;这是该技术的根本局限性。它们歪曲 RAG 结果的可能性,和它们歪曲其他训练数据的可能性一样大。这意味着它绝不能将其结果呈现为权威结论。

If you ask an AI to do research queries, every result should be presented as a list of citations. Moreover, the presentation should display each citation as a large object of in its own right, with clearly identified metadata, including not just the site where it was found but its publication date and, if possible, the name of the author. The literal, unmodified quotation (not from RAG, not a summary: a quote). 如果你要求 AI 进行研究查询,每一个结果都应该以引用列表的形式呈现。此外,呈现方式应将每一条引用作为一个独立的大对象展示,并附带清晰的元数据,不仅包括发现该内容的网站,还应包括发布日期,如果可能的话,还要包括作者姓名。以及字面意义上、未经修改的原文引用(不是来自 RAG 的摘要,而是直接引用)。