Ember-1
Ember-1
Ember-1: half the tokens, same answers Ember-1:Token 减半,效果如初
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens. Built on Kimi K3, it learned to cut unnecessary reasoning while keeping the thinking that matters. We tested it on external benchmarks, in live customer A/B tests, and on our own coding and agent workloads, and quality held up in every setting. Available today, Ember-1 kicks off an ongoing series of specialized models by Fireworks, shaped by what developers want next. Ember is just the start of what you could build with the Fireworks Training platform.
Ember-1 是 Fireworks Research 推出的全新专用模型,它在保持 Kimi K3 同等质量的同时,Token 使用量减少了 40%。该模型基于 Kimi K3 构建,学会了在剔除冗余推理过程的同时,保留关键的思考逻辑。我们通过外部基准测试、实时客户 A/B 测试以及我们内部的编程和智能体工作负载对其进行了验证,结果显示其在各种场景下均保持了高质量。Ember-1 即日起正式发布,它开启了 Fireworks 专用模型系列的新篇章,该系列将根据开发者的需求持续迭代。Ember 仅仅是利用 Fireworks 训练平台所能构建成果的开端。
How Fireworks Research built Ember-1 Fireworks Research 如何构建 Ember-1
We heard from users that they needed K3’s coding capabilities at a lower cost, because its long reasoning traces made automated coding expensive at scale. Turning down K3’s reasoning effort didn’t solve this. Lower effort settings gave up too much quality. To keep the quality and cut the tokens, the model had to learn to reason more efficiently, and that meant training it.
我们从用户那里了解到,他们需要 K3 的编程能力,但希望成本更低,因为其冗长的推理轨迹使得大规模自动化编程变得昂贵。调低 K3 的推理强度并不能解决这个问题,因为较低的强度设置会导致质量大幅下降。为了在保持质量的同时减少 Token,模型必须学会更高效地推理,这意味着必须对其进行针对性训练。
Getting there took serious research. Our team ran more than 50 training experiments and over 200 evaluations, and developed new training algorithms along the way to shorten reasoning without losing accuracy. We did it all on Fireworks Serverless Training. Because we didn’t have to provision or manage GPUs, we could launch experiments as soon as we had an idea, pay only for what we ran, and move from research to launch in a fraction of the usual time and cost.
实现这一目标需要深入的研究。我们的团队进行了 50 多次训练实验和 200 多次评估,并在过程中开发了新的训练算法,以在不损失准确性的前提下缩短推理过程。这一切都是在 Fireworks Serverless Training 上完成的。由于无需配置或管理 GPU,我们可以在产生想法后立即启动实验,仅为实际运行的部分付费,从而以极短的时间和极低的成本完成了从研究到发布的转化。
The problem: thinking models think too much 问题所在:思考型模型“想得太多”
Reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning rather than the answer itself. This thinking structure is expensive on a single request, but it gets much worse in multi-turn agentic workloads. Every turn replays all prior reasoning back to the model, so context grows roughly quadratically with the number of turns. Long reasoning traces from early turns get re-read (and re-billed) on every subsequent call.
像 Kimi K3 这样的推理模型,其生成的 Token 中大部分(有时超过 90%)用于内部推理,而非答案本身。这种思考结构在单次请求中成本高昂,而在多轮智能体工作负载中情况则更为严重。每一轮对话都会将之前所有的推理过程重新输入给模型,因此上下文长度随轮数呈二次方增长。早期轮次产生的冗长推理轨迹在后续的每一次调用中都会被重新读取(并重新计费)。
Is all that reasoning actually necessary? Our experiments said no. The reasoning Kimi K3 emits is far longer than the task requires, and the excess can be removed without touching the answer. This was how we created Ember-1, an economical version of Kimi K3 built from specialized intelligence.
所有这些推理真的有必要吗?我们的实验结果是否定的。Kimi K3 输出的推理过程远超任务所需,多余的部分可以在不影响答案的情况下被剔除。这就是我们如何创造出 Ember-1 的过程——一个基于专用智能构建的、经济高效版的 Kimi K3。
From an observation to a premium model 从观察到优质模型
Not all of K3’s reasoning is wasted. Some of it is self-reflection: revisiting an assumption, responding to feedback, or tracing an outcome back to an earlier decision can help the model recover from mistakes. The opportunity is to preserve this ability while reducing unnecessary reasoning and escaping unproductive loops. We believe that learning from tasks and environment feedback can teach the model to reason more efficiently while maintaining its capabilities.
并非 K3 的所有推理都是浪费。其中一部分是自我反思:重新审视假设、响应反馈或将结果追溯到之前的决策,这些都有助于模型从错误中恢复。我们的机会在于保留这种能力,同时减少不必要的推理并跳出低效的循环。我们相信,通过从任务和环境反馈中学习,可以教会模型在保持能力的同时进行更高效的推理。
The Specialized Intelligence Index: Ember-1 sets a Pareto frontier for Bedside Bench 专用智能指数:Ember-1 在 Bedside Bench 上树立了帕累托前沿
Earlier this week, we introduced the Specialized Intelligence Index (SII) to benchmark open, closed, and specialized models against real-world tasks created by industry experts. We evaluated Ember-1 on Doximity’s Bedside Bench, a physician-validated benchmark spanning 500 clinical cases across 10 specialized categories. The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.
本周早些时候,我们推出了专用智能指数 (SII),旨在针对行业专家创建的真实世界任务,对开源、闭源和专用模型进行基准测试。我们在 Doximity 的 Bedside Bench 上评估了 Ember-1,这是一个由医生验证的基准测试,涵盖了 10 个专业类别的 500 个临床案例。结果如何?Ember-1 在成本/任务维度上,为 Bedside Bench 树立了涵盖开源和闭源模型(包括 GPT-5.6 Sol、GPT-6 Astra 和 Claude Opus 5)的全新帕累托前沿。