Cline in Production: BYO-Key Costs, MCP Limits, and Terminal-Bench Results

Cline in Production: BYO-Key Costs, MCP Limits, and Terminal-Bench Results

Cline 在生产环境中的应用:自带密钥 (BYO-Key) 成本、MCP 限制与 Terminal-Bench 测试结果

Why we evaluated Cline: We did not test Cline because we needed another autocomplete tool. We tested it because flat per-seat AI IDE pricing makes cost attribution difficult, while closed agent runtimes make model and tool migrations expensive. 我们评估 Cline 的原因:我们测试 Cline 并非因为需要另一个自动补全工具。我们测试它是由于固定的 AI IDE 按席位定价使得成本归因变得困难,而封闭的智能体运行时(Agent Runtime)又导致模型和工具的迁移成本高昂。

Cline + Kimi K3 on Terminal-Bench 2.1: score up, spend down Cline + Kimi K3 在 Terminal-Bench 2.1 上的表现:分数提升,支出下降

  • Baseline pass rate (69/89) 77.5%/100
  • Confirmation pass rate (79/89) 88.8%/100
  • Baseline run cost $79
  • Confirmation run cost $49.80
  • 基准通过率 (69/89) 77.5%/100
  • 确认通过率 (79/89) 88.8%/100
  • 基准运行成本 $79
  • 确认运行成本 $49.80

The higher-scoring run on the same 89-task suite was also the cheaper one, because fewer sessions died in retries, loops, and self-termination. 在同一套 89 个任务的测试中,得分更高的那次运行反而成本更低,这是因为在重试、循环和自动终止过程中死掉的会话更少。

Cline separates the agent runtime from the inference provider. The VS Code extension supplies the agent loop, file operations, terminal integration, approvals, context management, and MCP client. We supply the model account and pay the inference provider directly. That separation gives us three things we cannot assume from a bundled IDE subscription: Cline 将智能体运行时与推理提供商分离开来。VS Code 插件负责提供智能体循环、文件操作、终端集成、审批、上下文管理以及 MCP 客户端。我们则提供模型账户并直接向推理提供商付费。这种分离为我们带来了捆绑式 IDE 订阅所无法保证的三点优势:

  1. We can select or replace the model independently of the editor.
  2. We can inspect token usage and assign inference cost to a task.
  3. We can expose internal services through MCP without first waiting for an IDE vendor to add a native integration.
  4. 我们可以独立于编辑器选择或更换模型。
  5. 我们可以检查 Token 使用情况,并将推理成本分配给特定任务。
  6. 我们可以通过 MCP 暴露内部服务,而无需等待 IDE 供应商添加原生集成。

That does not make Cline free. It exchanges predictable seat pricing for variable inference spend, provider rate limits, API-key management, and substantially more operational responsibility. 这并不意味着 Cline 是免费的。它用可预测的席位定价换取了可变的推理支出、提供商速率限制、API 密钥管理以及更多的运维责任。

We evaluated four practical questions: Can a developer install Cline without adopting another IDE? Can we switch between Anthropic and OpenAI-style providers without rewriting the workflow? Does MCP provide a usable portability boundary? Do the published Terminal-Bench numbers survive scrutiny as production evidence? 我们评估了四个实际问题:开发者能否在不更换 IDE 的情况下安装 Cline?我们能否在不重写工作流的情况下在 Anthropic 和 OpenAI 风格的提供商之间切换?MCP 是否提供了可用的可移植性边界?公布的 Terminal-Bench 数据能否经得起作为生产环境证据的审查?

The short answer is mixed. BYO-key model choice and usage-based accounting provide useful control, but we did not verify the extension’s installation or provider-configuration workflow. Our separate Python MCP SDK test failed during initialization; it does not establish a Cline client defect. The benchmark is much more useful as a harness-debugging case study than as a definitive “Cline beats Cursor” ranking. 简短的回答是:情况不一。自带密钥(BYO-key)的模型选择和基于用量的核算提供了有效的控制,但我们并未验证该插件的安装或提供商配置工作流。我们独立的 Python MCP SDK 测试在初始化期间失败了;但这并不代表 Cline 客户端存在缺陷。该基准测试作为测试框架调试的案例研究,比作为“Cline 击败 Cursor”的定论排名更有价值。

Cline’s SDK production architecture also matters. We can instrument model-call, tool-use, session-end, and usage events, including token counts and finish reasons. The SDK supports iteration limits, per-turn token limits, loop detection, mistake limits, and host-side cancellation. Those are the controls we expect from an agent runtime. They are not proof that a developer workstation is safely sandboxed. Our evaluation therefore treated Cline as an agent harness with an editor front end, not as a security boundary. Cline 的 SDK 生产架构也很重要。我们可以对模型调用、工具使用、会话结束和使用事件进行监测,包括 Token 计数和完成原因。该 SDK 支持迭代限制、单轮 Token 限制、循环检测、错误限制以及主机端取消。这些是我们期望从智能体运行时中获得的控制手段。它们并不能证明开发工作站已安全沙箱化。因此,我们的评估将 Cline 视为带有编辑器前端的智能体测试框架,而非安全边界。

Installing Cline, configuring MCP, and running Terminal-Bench

安装 Cline、配置 MCP 并运行 Terminal-Bench

Installing and verifying the extension: For a team rollout, we would verify the current Marketplace extension identifier and test installation in a clean VS Code profile. The commands below are illustrative; we have not verified their identifier or execution: 安装并验证插件:对于团队推广,我们会验证当前的 Marketplace 插件标识符,并在干净的 VS Code 配置文件中测试安装。以下命令仅供参考;我们尚未验证其标识符或执行情况:

code --install-extension saoudrizwan.claude-dev
code --list-extensions --show-versions | grep -i saoudrizwan.claude-dev

We would confirm the publisher, display name, and exact extension identifier against the current Marketplace listing before using these commands. For a team rollout, we would pin and test a known version before broad deployment: 在使用这些命令之前,我们会根据当前的 Marketplace 列表确认发布者、显示名称和确切的插件标识符。对于团队推广,我们会在大规模部署前锁定并测试一个已知版本:

code --install-extension saoudrizwan.claude-dev@<approved-version> --force

Cline supports BYO-key providers, including Anthropic and OpenAI. We would verify the extension’s current provider interface and credential-storage behavior before rollout, and keep keys out of workspace settings and committed files. For Anthropic or OpenAI, we would confirm the current provider-specific setup instructions, configure a dedicated development key, and select the intended model. We have not verified the extension’s exact settings labels, compatible-endpoint fields, or approval controls. Our rollout plan would begin with a read-only task before enabling terminal commands or file writes. We would use separate development and production provider keys and configure provider-side budget alerts. A BYO-key extension installed on every laptop creates a larger credential-management surface than a centrally administered seat license. Cline 支持自带密钥(BYO-key)的提供商,包括 Anthropic 和 OpenAI。在推广前,我们会验证插件当前的提供商接口和凭据存储行为,并将密钥排除在工作区设置和已提交的文件之外。对于 Anthropic 或 OpenAI,我们会确认当前特定于提供商的设置说明,配置专用的开发密钥,并选择目标模型。我们尚未验证该插件确切的设置标签、兼容端点字段或审批控制。我们的推广计划将从只读任务开始,然后再启用终端命令或文件写入。我们会使用独立的开发和生产提供商密钥,并配置提供商侧的预算警报。与集中管理的席位许可相比,在每台笔记本电脑上安装自带密钥的插件会产生更大的凭据管理面。

Wiring a local MCP stub

连接本地 MCP 存根 (Stub)

The following is an illustrative Python MCP server using the SDK version pinned for our local test attempt. This example was not successfully validated, and it is not a verified Cline integration: 以下是一个示例性的 Python MCP 服务器,使用了我们本地测试尝试时锁定的 SDK 版本。此示例未成功验证,也不是经过验证的 Cline 集成:

# server.py
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("effloow-local-stub")

@mcp.tool()
def add_ticket(label: str, priority: int) -> dict:
    """Create a local test ticket without external side effects."""
    if priority < 1 or priority > 5:
        raise ValueError("priority must be between 1 and 5")
    return {
        "accepted": True,
        "ticket": {
            "label": label,
            "priority": priority,
        },
    }

if __name__ == "__main__":
    mcp.run(transport="stdio")

We installed its dependencies in an isolated environment: 我们在隔离环境中安装了其依赖项:

python -m venv .venv
. .venv/bin/activate
pip install "mcp==1.9.4" "pydantic==2.11.5" "anyio==4.9.0"

The following is an unverified registration example, not a configuration we tested in Cline. We would confirm the current client schema, configuration location, and approval fields before using it: 以下是一个未经证实的注册示例,并非我们在 Cline 中测试过的配置。在使用它之前,我们会确认当前的客户端架构、配置位置和审批字段:

{
  "mcpServers": {
    "effloow-local-stub": {
      "command": "/absolute/path/to/project/.venv/bin/python",
      "args": ["/absolute/path/to/project/server.py"],
      "disabled": false,
      "autoApprove": []
    }
  }
}

For a future integration test, we would use absolute interpreter and server paths and check the environment supplied by the client. We would also verify the meaning of approval fields rather than assume that an empty autoApprove list enforces the intended policy. State-changing tools should require explicit approval. Our isolated Python MCP SDK handshake did not complete successfully. The test failed during initialization with McpError: Connection closed and exit code 1, producing no tool observations. Dependency setup completed, but we did not verify Cline’s registration structure. 对于未来的集成测试,我们将使用绝对解释器和服务器路径,并检查客户端提供的环境。我们还会验证审批字段的含义,而不是假设空的 autoApprove 列表就能强制执行预期的策略。改变状态的工具应要求明确的审批。我们隔离的 Python MCP SDK 握手未能成功完成。测试在初始化期间失败,报错 McpError: Connection closed 和退出代码 1,未产生任何工具观测结果。依赖项设置已完成,但我们尚未验证 Cline 的注册结构。