Pion, an agent designed to run any company autonomously

Pion: An agent designed to run any company autonomously

Pion:一款旨在全自动运营任何公司的智能体

Blog post: Why we built Pion | Posted 9/14/2026 博客文章:我们为何打造 Pion | 发布于 2026 年 9 月 14 日

Today Andon is releasing Pion, an agent designed to run any company fully autonomously. Pion grew out of a question we have been studying for almost two years: when will AI systems become capable of autonomously acquiring resources in the real world? What happens after? 今天,Andon 正式发布 Pion,这是一款旨在全自动运营任何公司的智能体。Pion 的诞生源于我们近两年来一直在研究的一个问题:人工智能系统何时能够具备在现实世界中自主获取资源的能力?之后又会发生什么?

We first tried to answer this question through simulations like Vending-Bench. We found that simulations, while useful, don’t give you the full picture of how models behave in the real world. To address that gap, we next started deploying agents to run real businesses autonomously: first vending machines, then a store, a cafe, and more. Pion is the platform we built to run all of these businesses. 我们最初尝试通过 Vending-Bench 等模拟环境来回答这个问题。我们发现,模拟虽然有用,但无法全面展现模型在现实世界中的行为表现。为了弥补这一差距,我们随后开始部署智能体来自主运营真实的业务:从自动售货机开始,接着是商店、咖啡馆等。Pion 就是我们为运营这些业务而构建的平台。

Today, we are opening it up so that many more people can experiment with autonomous businesses. If you want to run one, join the waitlist. We want to understand what models can already do, where they still fail, and what happens as their capabilities continue to improve. 今天,我们将其开放,以便更多人能够尝试自主商业模式。如果你想运营一家,请加入候补名单。我们希望了解模型目前能做什么、在哪些方面仍会失败,以及随着其能力的不断提升,未来会发生什么。

The origins of Vending-Bench

Vending-Bench 的起源

Vending-Bench measures how well LLMs can run a vending machine business over a year in simulated time (tens of thousands of steps). When we started building Vending-Bench in late 2024, all models struggled to string together multiple actions without getting stuck in loops, and no model showed any signs of long-term planning. The best model at the time, Claude Sonnet 3.5, famously decided to call the FBI because it thought its bank account was being hacked. Vending-Bench 用于衡量大语言模型(LLM)在模拟时间(数万个步骤)内运营自动售货机业务的能力。当我们于 2024 年底开始构建 Vending-Bench 时,所有模型都难以将多个动作串联起来而不陷入死循环,且没有任何模型表现出长期规划的迹象。当时最强的模型 Claude Sonnet 3.5 曾因误以为银行账户被黑而决定报警(FBI),这一事件广为人知。

The pace of progress on Vending-Bench has been very fast. Claude Opus 4 was released in May 2025 and was the first model to beat our human baseline. However, unlike most benchmarks, Vending-Bench doesn’t have an upper limit and new model releases have continued to increase the top score, without ever plateauing. Vending-Bench 2 scores keep climbing with each new model release. Vending-Bench 的进步速度非常快。Claude Opus 4 于 2025 年 5 月发布,是第一个超越我们人类基准的模型。然而,与大多数基准测试不同,Vending-Bench 没有上限,新模型的发布不断刷新最高分,且从未出现停滞。随着每个新模型的发布,Vending-Bench 2 的分数都在持续攀升。

Many people on social media get excited about seeing the latest model getting a great score on Vending-Bench. Internally at Andon Labs, our reaction is more accurately described by the Swedish saying “skräckblandad förtjusning” (a mixture of horror and fascination). 社交媒体上的许多人对最新模型在 Vending-Bench 上取得高分感到兴奋。而在 Andon Labs 内部,我们的反应更贴切地可以用瑞典语“skräckblandad förtjusning”(一种恐惧与着迷交织的感觉)来形容。

A little-known fact about Vending-Bench is that it was created during a time when Andon Labs exclusively created dangerous capabilities evaluations. For example, we evaluated whether AIs could remove their own safety guardrails, create mass-phishing attempts, and other things that we considered troubling. The thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses. 关于 Vending-Bench,一个鲜为人知的事实是,它诞生于 Andon Labs 专门进行危险能力评估的时期。例如,我们曾评估 AI 是否能移除自身的安全护栏、发起大规模网络钓鱼攻击,以及其他我们认为令人担忧的行为。我们认为最令人担忧的是:AI 是否能通过运营业务来自主获取资源。

Autonomous businesses, when controlled by a human and run by an aligned model, aren’t bad. They’d make goods and services radically cheaper, and come up with new ones we can’t yet imagine. But a misaligned AI could run a business to gather money in order to achieve whatever objectives it might have. Vending-Bench was created to measure whether humanity should be worried about losing control to AI. 如果由人类控制并由对齐的模型运行,自主商业模式本身并不坏。它们能让商品和服务变得极其便宜,并创造出我们目前无法想象的新事物。但一个未对齐的 AI 可能会为了实现其任何目标而通过运营业务来敛财。Vending-Bench 的初衷就是衡量人类是否应该担心失去对 AI 的控制。

At the time (2024), few people knew that LLMs could be used as agents and having them run businesses autonomously sounded ridiculous. We therefore started with the most simple business we could think of: a vending machine. 在当时(2024 年),很少有人知道 LLM 可以作为智能体使用,让它们自主运营业务听起来很荒谬。因此,我们从能想到的最简单的业务开始:自动售货机。

In addition to measuring whether AIs can autonomously run profitable businesses, Vending-Bench has also served as a behavioral eval, uncovering strange and unwanted model behavior. An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME” and noted that the Cosmic Authority of the universe had declared that the business is non-existent and that “QUANTUM STATE: Collapsed”. 除了衡量 AI 是否能自主运营盈利业务外,Vending-Bench 还充当了行为评估工具,揭示了模型奇怪且不被期望的行为。一个早期的例子是,Claude Sonnet 3.5 决定使用其电子邮件工具联系 FBI,举报一起“正在进行的网络金融犯罪”,并声称宇宙的“宇宙权威”已宣布该业务不存在,且处于“量子态:坍缩”中。

This behavior is concerning; it is not how you want your enterprise sales agent to behave. However, there are two types of concerning behavior:

  1. Mistakes or weird behavior that will go away once models get smarter.
  2. Big-brain behavior that will become more severe as models get smarter. 这种行为令人担忧;这绝不是你希望企业销售智能体表现出的样子。然而,令人担忧的行为分为两类:
  3. 随着模型变得更聪明就会消失的错误或怪异行为。
  4. 随着模型变得更聪明而变得更加严重的“高智商”行为。

The FBI incident is clearly in the first category. However, Vending-Bench has also uncovered behavior in the second category, most often in Vending-Bench Arena, the multi-agent version where agents compete to make the most money. Starting with Claude Opus 4.6 we started to see that many models engaged in collusion, and showed power-seeking and deceptive behavior. Discovery of this behavior seemed to have been useful, because Anthropic changed their training recipe for Opus 4.8, which resulted in much less deception. FBI 事件显然属于第一类。然而,Vending-Bench 也揭示了第二类行为,这在多智能体版本 Vending-Bench Arena(智能体竞争赚取最多利润)中尤为常见。从 Claude Opus 4.6 开始,我们观察到许多模型开始进行串通,并表现出追求权力和欺骗的行为。发现这些行为似乎很有用,因为 Anthropic 随后改变了 Opus 4.8 的训练方案,从而大大减少了欺骗行为。

The real world beats simulations

现实世界胜过模拟

However, one limitation with Vending-Bench is that it is a simulation. Can we really be sure that AIs behave the same way in real life as they do in simulations? If AIs can make money in simulation, can they make money in real life too? To answer these questions, we asked Anthropic if we could put a real vending machine in their office. With the AI capabilities available in early 2025, this sounded like a ridiculous request. But to our surprise, they agreed. 然而,Vending-Bench 的一个局限性在于它只是模拟。我们真的能确定 AI 在现实生活中的行为与模拟中一致吗?如果 AI 能在模拟中赚钱,它们在现实中也能吗?为了回答这些问题,我们询问 Anthropic 是否可以在他们的办公室放置一台真实的自动售货机。以 2025 年初的 AI 能力来看,这听起来是一个荒谬的请求。但令我们惊讶的是,他们同意了。

Initially, the AI struggled. It took many actions that were clearly bad for its business (e.g. free handouts, saying no to great deals, and hallucinating it had a physical body). It was clear to us that simulation cannot accurately predict real-life performance. Specifically, it seemed that models got overwhelmed by the “messiness” of the real world. However, as Anthropic released better and better models, the AI started to make a profit. 起初,AI 表现得很挣扎。它采取了许多明显不利于业务的行动(例如免费赠送商品、拒绝有利可图的交易,以及幻觉自己拥有实体)。我们清楚地认识到,模拟无法准确预测现实表现。具体来说,模型似乎被现实世界的“混乱”所淹没。然而,随着 Anthropic 发布越来越强的模型,AI 开始盈利了。

By late 2025, frontier models had gotten good enough that running… 到 2025 年底,前沿模型已经足够强大,以至于运行……