Kimi K3-256k

Kimi K3-256k

Model Configuration

This page covers the models Kimi Code provides and how to switch between them in each client. 模型配置 本页面介绍了 Kimi Code 提供的模型以及如何在各个客户端中进行切换。

Model Overview

Kimi Code currently offers two models—Kimi K3 and Kimi K2.7 Code—across four model IDs, selectable by model ID in clients or third-party tools. 模型概览 Kimi Code 目前提供两款模型——Kimi K3 和 Kimi K2.7 Code,共计四个模型 ID,可在客户端或第三方工具中通过模型 ID 进行选择。

Model specs:

Recommended model launch k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k. Ideal for everyday Q&A, code completion, routine feature development, and single-file or small-file edits — video input is not supported. 模型规格: 推荐模型 k3-256k 现已上线。在 256k 上下文内,它提供与 k3 (1M) 相同的效果。k3 (1M) 的配额消耗约为 k3-256k 的两倍。它非常适合日常问答、代码补全、常规功能开发以及单文件或小文件编辑——不支持视频输入。

💓 Reminder

Switching from K3 (1M) to K3-256k: When switching from k3 (1M) to k3-256k, if the current session’s context already exceeds 256k, some coding tools such as Kimi Code CLI and Claude Code will perform a compact on the tool side. 💓 提醒 从 K3 (1M) 切换到 K3-256k: 当从 k3 (1M) 切换到 k3-256k 时,如果当前会话的上下文已超过 256k,Kimi Code CLI 和 Claude Code 等部分编码工具会在工具端执行压缩(compact)。

Switching recommendations: (1) Because different agent tools handle this differently, manually run compact once before switching to compress the context to within 256k. This preserves the key points of the task, keeps the session intact, and lets you benefit from more durable quota after switching. (2) If the conversation history includes video files, switching directly will fail because K3-256k does not support video input. Please compact first, then switch. 切换建议: (1) 由于不同的 Agent 工具处理方式不同,建议在切换前手动运行一次 compact,将上下文压缩至 256k 以内。这能保留任务要点,保持会话完整,并让你在切换后享受更持久的配额。(2) 如果对话历史包含视频文件,直接切换会失败,因为 K3-256k 不支持视频输入。请先压缩,再切换。

Switching from K3-256k to K3 (1M): When switching from k3-256k to k3 (1M), if k3-256k is close to the 256k limit and you don’t want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache. 从 K3-256k 切换到 K3 (1M): 当从 k3-256k 切换到 k3 (1M) 时,如果 k3-256k 已接近 256k 限制,且你不希望通过压缩丢失信息,可以直接切换到 1M。当前版本从 256k 切换到 1M 不会影响缓存。


Model ID Table (模型 ID 表)

Model IDModel versionDescriptionSpeedContext windowReasoningThinkingAvailabilityMultimodal input
k3Kimi K3Kimi’s most capable flagship coding model: 2.8T parameters, 1M context windowRegularUp to 1Mlow/high/maxONModerato+Image, video
k3-256kKimi K3The 256K context version of Kimi K3, effectively reducing consumptionRegular256k onlylow/high/maxONModerato+Image only
kimi-for-codingKimi K2.7 CodeGood at code completion and routine development tasksRegular256k--All membersImage, video
kimi-for-coding-highspeedK2.7 Code HighSpeedThe high-speed version of K2.7 Code, with the same coding ability and ~5–6× faster outputHighSpeed256k--Allegretto+Image, video

Why did usage go up after the new model launched?

After switching models, the context cache built earlier no longer hits on the new model, so that context has to be re-prefilled. Usage therefore looks higher right after switching. Recommended action: Start a new session when using the new model: this gives better results and lower consumption. 为什么新模型发布后使用量上升了? 切换模型后,之前构建的上下文缓存无法在新模型上命中,因此上下文必须重新预填充(re-prefilled)。因此,切换后立即显示的使用量会更高。建议操作:使用新模型时开启新会话,这样效果更好且消耗更低。

Why do I still get a 401 with the correct model ID?

When the requested capability exceeds your plan’s entitlements, the server returns 401. Three common cases:

  1. No K3 access: your plan is below Moderato and can’t call k3, k3-256k — upgrade to Moderato or above.
  2. No 1M access: on a Moderato plan, k3 supports up to 256K context; up to 1M context is available on Allegretto and higher tiers. k3-256k has a fixed 256K context limit.
  3. No HighSpeed access: some plans don’t include HighSpeed — upgrade to Allegretto or a higher tier to call kimi-for-coding-highspeed. 为什么模型 ID 正确却依然收到 401 错误? 当请求的能力超出你的套餐权限时,服务器会返回 401。常见的三种情况:
  4. 无 K3 访问权限: 你的套餐低于 Moderato,无法调用 k3 或 k3-256k —— 请升级至 Moderato 或更高版本。
  5. 无 1M 访问权限: 在 Moderato 套餐中,k3 支持最高 256K 上下文;最高 1M 上下文仅在 Allegretto 及更高层级提供。k3-256k 有固定的 256K 上下文限制。
  6. 无 HighSpeed 访问权限: 部分套餐不包含 HighSpeed —— 请升级至 Allegretto 或更高层级以调用 kimi-for-coding-highspeed。

How to Switch Models

Usage notes:

  • Start a new session when switching model IDs: switching models invalidates the context cache you’ve built up.
  • Fill in the Model ID, not the model version name: use IDs like k3, k3-256k, etc.
  • K3 / K2.7 without Thinking routes to K2.6: keep Thinking on to use K3 or K2.7 Code. 如何切换模型 使用注意事项:
  • 切换模型 ID 时开启新会话: 切换模型会使之前构建的上下文缓存失效。
  • 填写模型 ID,而非模型版本名称: 请使用 k3k3-256k 等 ID。
  • 未开启 Thinking 的 K3 / K2.7 会路由至 K2.6: 保持 Thinking 开启以使用 K3 或 K2.7 Code。