Your agents need a desk, not another chat pane
Your agents need a desk, not another chat pane
你的智能体需要的是一张办公桌,而不是另一个聊天窗口
Most agent demos still look the same. You type into a box. Something happens somewhere else. A wall of tool calls scrolls by. Then a summary appears and you are supposed to trust it. That model works for narrow tasks. It falls apart the moment the agent has to use the same computer you use: browsers, sheets, tickets, messy UIs that were never designed for automation. If you cannot see the pointer, you are not collaborating. You are waiting. 大多数智能体演示看起来都大同小异:你在对话框里输入指令,别处发生了一些变化,一连串工具调用记录飞速滚动,最后弹出一个总结,而你只能选择相信它。这种模式在处理单一任务时尚可,但一旦智能体需要操作你日常使用的电脑——比如浏览器、表格、工单系统,以及那些从未为自动化设计过的混乱界面时,它就彻底失效了。如果你看不见光标,那就不叫协作,那叫等待。
I build product and engineering for an infinite canvas OS. The design bet we keep returning to is simple: agents and humans should share a spatial workspace. Same board. Separate cursors. Interruptible work. Here is what that actually implies when you try to ship it. 我正在为一款“无限画布”操作系统负责产品与工程。我们始终坚持的设计理念很简单:智能体与人类应该共享一个空间化的工作区。同一块画板,独立的光标,可随时中断的工作流。当你真正尝试实现它时,这意味着以下几点。
Chat is a log. A desk is an address space. Chat is great for intent. “Book the cabin for Oct 8–13 if it is under $250.” Fine. What chat is bad at is state you can point at. Developers already know this from debugging. A stack trace is not the same as an open debugger with the heap still live. Agents that only report back in text force you to reconstruct the world from narration. A spatial workspace flips that. Windows sit next to each other. The agent’s browser tab is a place, not a sentence. Your notes stay beside the sheet it is filling. When something looks wrong, you do not ask for another summary. You look. That is closer to how humans already work together on a real desk. One person owns the left side of the table. Another owns the right. Nobody waits for a paste into Slack. 聊天是日志,而办公桌是地址空间。聊天非常适合表达意图,比如“如果价格低于250美元,就预订10月8日至13日的客舱”,这没问题。但聊天不擅长处理那些你可以直接指出的状态。开发者在调试时深知这一点:堆栈跟踪(stack trace)与实时运行的调试器完全不同。只通过文本汇报的智能体,强迫你通过叙述来重构世界;而空间化工作区则颠覆了这一点。窗口并排摆放,智能体的浏览器标签页是一个“地点”,而不是一句话。你的笔记就放在它正在填写的表格旁边。当发现不对劲时,你不需要索要总结,你只需要看一眼。这更接近人类在真实办公桌上的协作方式:一个人负责桌子左侧,另一个人负责右侧,没人需要等待对方把内容粘贴到 Slack 里。
Agents that steal your mouse are not teammates. On a normal desktop there is one pointer. If an agent uses the computer through that pointer, two bad outcomes show up immediately. Either the agent takes your mouse and you sit there while a spinner owns the session. Or the agent runs out of sight and you only see the after-action report. Neither feels like pair programming. Pair programming works because you can watch the other person’s hands. So give the agent its own cursor. Label it. Color it. Keep yours. Now you can keep typing in a sheet while the agent works in a booking page next door. Collision becomes a layout problem, not a trust problem. We ended up treating each agent seat as a mouse and keyboard inside the compositor. Add an agent, drop a new cursor into a new window. Up to sixteen on one canvas in our case. Helpers get attached to a lead, one task per window, so the board does not turn into a pile of anonymous activity. 抢夺你鼠标的智能体不是队友。在普通桌面上只有一个光标,如果智能体通过这个光标操作电脑,会立即出现两个糟糕的结果:要么智能体夺走鼠标,你只能干坐着看加载转圈;要么智能体在后台运行,你只能看到事后的报告。这两种感觉都不像结对编程。结对编程之所以有效,是因为你能看到对方的手。所以,给智能体一个独立的光标,给它贴上标签,标上颜色,同时保留你自己的光标。现在,你可以在表格里继续打字,而智能体在旁边的预订页面工作。冲突变成了布局问题,而不是信任问题。我们最终将每个智能体席位视为合成器内的一个鼠标和键盘。每增加一个智能体,就在新窗口中投放一个新光标。在我们的系统中,一个画布上最多可以容纳16个。辅助智能体挂载在主智能体下,每个窗口处理一个任务,这样画板就不会变成一堆杂乱无章的活动。
Interruptibility is the feature people actually want. The first time someone watches an agent click the wrong listing, they do not ask for a better model. They ask for a stop button that works mid-click. Pause. Take the window back. Type a correction without restarting the whole run. Those controls sound small. They are the difference between “autonomy” and “supervision.” Autonomy without interruption is just a long-running script with better prose. A useful pattern: the lead agent writes a plan you can edit before it goes. Helpers execute pieces in view. A separate check reads a fresh screenshot before anything is marked done. “Done” stops meaning “the model said it finished.” If you are designing an agent OS or agentic OS layer, bake interruptibility in early. Retrofitting it onto a fire-and-forget runner is painful. “可中断性”才是人们真正想要的功能。当用户第一次看到智能体点错了列表时,他们不会要求换一个更好的模型,他们想要的是一个在点击过程中就能生效的“停止”按钮。暂停,接管窗口,输入修正,而无需重启整个流程。这些控制功能听起来很小,但它们是“自主”与“监管”的区别。没有中断能力的自主,不过是一个文笔更好的长脚本。一个实用的模式是:主智能体写好计划,你在执行前进行编辑;辅助智能体在视野内执行任务片段;在标记为“完成”之前,通过独立的检查机制读取最新的截图。这样,“完成”就不再仅仅意味着“模型说它做完了”。如果你正在设计智能体操作系统或智能体层,请尽早内置可中断性。在“发射后不管”的运行器上后期加装这个功能是非常痛苦的。
Shared canvas beats shared screen for multiplayer work. Developers already hate the remote debugging ritual. One person shares. Everyone else narrates. “Scroll up.” “No, the other tab.” “Wait, who has control?” Screen share is a video of a desk. A shared canvas is the desk. On a shared spatial workspace, every cursor works. A teammate can type into a window that still runs on your machine. Agents belonging to a teammate show up labeled as theirs. The room stays when the call ends. That last part matters more than it sounds. Most collaborative sessions die with the Zoom. The board evaporates. Tomorrow you rebuild context from memory and a half-updated doc. Saved layout is underrated infrastructure. Apps, positions, tabs, still there in the morning. Boring. Extremely useful. 对于多人协作,共享画布胜过共享屏幕。开发者早已厌倦了远程调试的仪式:一个人共享屏幕,其他人负责指挥——“往上滚”、“不,是另一个标签页”、“等等,谁在控制?”。屏幕共享只是桌面的视频,而共享画布本身就是那张桌子。在共享的空间工作区中,每个光标都能工作。队友可以在你机器上运行的窗口中输入内容。属于队友的智能体会被标记出来。通话结束后,房间依然存在。最后这一点比听起来更重要。大多数协作会话随着 Zoom 会议结束而消亡,画板随之蒸发。第二天,你只能靠记忆和一份半更新的文档来重建上下文。保存布局是一种被低估的基础设施。应用、位置、标签页,第二天早上依然在那里。这很无聊,但极其有用。
What I would require before calling something “human + agent workspace”: 在称呼某样东西为“人机协作工作区”之前,我会有以下要求:
- Visible agent input (cursor, rings, typing), not only a transcript
- 可见的智能体输入(光标、光圈、打字过程),而不仅仅是文字记录
- Agent windows that land beside yours instead of covering them
- 智能体窗口应并排显示,而不是覆盖你的窗口
- Pause / steer / reclaim without restarting the job
- 无需重启任务即可暂停、引导或收回控制权
- Verification that is not the same process that did the work
- 验证过程不能与执行任务的过程是同一个
- A place for your own notes and apps to remain in view while agents run
- 在智能体运行时,你的笔记和应用依然可见
- Multiplayer that does not serialize everyone onto one pointer
- 不会将所有人序列化到同一个光标上的多人协作模式
You can build pieces of this with existing tools. Browser automation plus screen share plus a chat sidebar gets you part of the way. The failure mode is the same every time: the human loses the spatial picture, and the agent becomes a black box again. We built Novastart around the opposite assumption — that the computer itself should be the canvas where people, apps, and agents stay in sight. Install as an app on Windows or Intel Mac to try it, or as an OS on a dedicated PC if you want the full speed. 你可以用现有工具构建其中的一部分。浏览器自动化加屏幕共享再加一个聊天侧边栏,可以实现部分功能。但失败模式总是相同的:人类失去了空间感,智能体再次变成了一个黑盒。我们构建 Novastart 的初衷恰恰相反——电脑本身应该成为一个画布,让人们、应用和智能体始终处于视野之内。你可以将其作为 Windows 或 Intel Mac 上的应用进行尝试,或者如果你想要极致速度,也可以将其作为操作系统安装在专用 PC 上。
Closing note for people shipping agent products: Stop asking whether users “trust AI.” Ask whether they can see what it is doing, stop it, and keep working in the same space. If your agent only lives in a pane, you have built a clever assistant. If it sits at the desk with its own hands, you have started building software people can actually co-operate with. That is the bar I care about. 给发布智能体产品的同行们的一点结语:别再问用户是否“信任 AI”了。问问他们是否能看到智能体在做什么,是否能随时叫停,并能在同一个空间里继续工作。如果你的智能体只活在一个窗口里,你只是造了一个聪明的助手;如果它坐在桌子旁,拥有自己的“双手”,你才真正开始构建人们能够与之协作的软件。这才是我在意的标准。