Private LLM Options: Local, Cloud, or Confidential?
Private LLM Options: Local, Cloud, or Confidential?
私有大模型选项:本地部署、云端服务,还是机密计算?
You want a private LLM, and the obvious move is to buy a box and run it yourself. Before you do, it’s worth understanding what that hardware really costs, what happens when it’s out of date in two years and you’re still paying it off, and what the other options are. There are three routes: your own hardware, a cloud provider under contract, or confidential inference on hardware that can’t read the prompt in the first place. This post walks through your options for Private LLMs, what each costs, and a deeper look into AI privacy when you’re sharing the agent with the rest of your team.
如果你想要一个私有大模型(LLM),最直观的做法就是买台机器自己运行。但在行动之前,你需要了解这些硬件的真实成本,以及当两年后硬件过时而你还在还贷时会发生什么,并看看还有哪些其他选择。目前有三条路径:使用自己的硬件、签署合同的云服务商,或者在无法读取提示词(Prompt)的硬件上进行机密推理。本文将带你了解私有大模型的选项、各自的成本,并深入探讨当你与团队共享 AI 代理时的隐私问题。
What does “private” mean for an LLM?
“私有”对于大模型意味着什么?
The term Private LLM / AI is used across various contexts. Physical privacy means the prompt never leaves a machine you control. Contractual privacy means it leaves, but the provider has agreed in writing not to train on it, share it, or keep it past a set window. Technical privacy means it leaves, but it’s processed inside hardware that the operator can’t see into, and you can check that yourself. Who is able to read the prompt, and is the thing stopping them a wall, a promise, or a chip?
“私有大模型/AI”这一术语在不同语境下有不同的含义。物理隐私意味着提示词永远不会离开你控制的机器。合同隐私意味着提示词虽然离开了你的机器,但服务商已书面承诺不会利用它进行训练、分享或在规定期限后保留。技术隐私意味着提示词虽然离开了你的机器,但它是在操作员无法窥视的硬件中处理的,且你可以自行验证这一点。究竟谁能读取提示词?阻挡他们的究竟是一堵墙、一份承诺,还是一块芯片?
Option 1: a local LLM on your own hardware
选项 1:在自有硬件上运行本地大模型
A local LLM is the cleanest answer to the privacy question. Nothing leaves the building, there’s no provider to trust, and no terms of service to reread every time they change. A current hardware roundup puts an RTX 5090 at $1,999 MSRP (USD) with 32 GB of VRAM, managing roughly 25 to 30 tokens per second on a 70B model, and a 128 GB Mac Studio at $3,500 to $4,000 for similar speeds (fungies.io hardware guide). That’s a perfectly good setup for one person.
本地大模型是解决隐私问题最彻底的方案。数据不出门,无需信任任何服务商,也不必担心服务条款变更。目前的硬件行情显示,RTX 5090 建议零售价为 1,999 美元(配备 32GB 显存),在运行 70B 模型时每秒可处理约 25 到 30 个 Token;而 128GB 内存的 Mac Studio 售价在 3,500 到 4,000 美元之间,性能相当(参考 fungies.io 硬件指南)。对于个人用户来说,这已经是非常好的配置了。
Tradeoffs are: Concurrency (one card serving one conversation), Model ceiling (open-weight models only), Flexibility (hardware becomes outdated), and Your time (maintenance). So local is the most private option and often the cheapest per token once the hardware is paid off. It is a large upfront cost and is likely to get outdated soon.
其权衡点在于:并发性(一张卡只能处理一个对话)、模型上限(仅限开源权重模型)、灵活性(硬件易过时)以及你的时间成本(维护工作)。因此,本地部署是最私密的选项,一旦硬件成本摊销完毕,其单 Token 成本通常也是最低的。但它的前期投入巨大,且很快就会面临过时。
Option 2: a cloud provider with a contract
选项 2:签署合同的云服务商
This is where most privacy-first enterprises land. The big three clouds all sell access to frontier models under enterprise terms that are a long way from pasting things into a consumer chat app. The privacy here is contractual, so the detail is in the documents. AWS Bedrock, Microsoft Azure, and Google Cloud all offer specific configurations to ensure data isn’t used for training or retained beyond a certain period. The provider’s staff and systems are technically able to see the prompt; and the thing stopping them is policy. For a lot of teams that’s fine. For source code, client documents or anything you’d lose sleep over, it’s worth knowing that’s the shape of the guarantee.
这是大多数注重隐私的企业最终选择的方案。三大云厂商都提供企业级的前沿模型访问权限,这与在消费级聊天应用中粘贴内容有着本质区别。这里的隐私是基于合同的,因此细节都在合同条款中。AWS Bedrock、Microsoft Azure 和 Google Cloud 都提供特定的配置,以确保数据不会被用于训练或在规定期限后被删除。从技术上讲,服务商的员工和系统是有能力看到提示词的,阻挡他们的仅仅是政策。对于许多团队来说这已经足够了。但对于源代码、客户文档或任何让你彻夜难眠的敏感信息,你需要清楚这种保障的本质。
Option 3: a private LLM gateway
选项 3:私有大模型网关
The third option is newer and sits between the other two. A private LLM gateway is an API you call like any cloud endpoint, except the gateway runs inside a trusted execution environment (TEE). The short version of a TEE is that the chip encrypts the memory of whatever is running inside it, so the host machine and the people operating it can’t read what’s in there. You rent the hardware like cloud, but the privacy is enforced by the chip, closer to local.
第三种选项较新,介于前两者之间。私有大模型网关是一个像普通云端接口一样调用的 API,但该网关运行在可信执行环境(TEE)中。简单来说,TEE 通过芯片加密运行中的内存,使得宿主机及其操作人员无法读取其中的内容。你像租用云服务一样租用硬件,但隐私是由芯片强制执行的,这更接近本地部署的安全性。
There are two different guarantees: Confidential models (open-weight models run inside the enclave, so neither the gateway nor the host can read the prompt) and Frontier models (the route is sealed, but the model provider like Anthropic or OpenAI may still see the prompts). The benefits of private LLM Gateways are: No upfront cost.
这里有两种不同的保障:机密模型(开源权重模型在安全区内运行,网关和宿主机都无法读取提示词)和前沿模型(传输路径是加密的,但 Anthropic 或 OpenAI 等模型提供商仍可能看到提示词)。私有大模型网关的优势在于:无需前期投入。