Jeeves. Reasoning improves Jev-like decision models
Jeeves: Reasoning improves Jev-like decision models
Jeeves: 推理能力提升了类 Jev 决策模型
A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO. Acknowledgements Inspired by Kev. 这是一个带有扩散草稿器(diffusion drafter)的 Jev 风格推理分类器,通过 SFT(监督微调)和 CISPO 进行训练。致谢:灵感源自 Kev。
Highlights
亮点
A 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, with a block-4 diffusion drafter and the full training code and train/dev/test data. 这是一个 9B 参数的类 Jev 模型(Qwen3.5-9B,使用 LoRA 和指针头),在决策前会进行思考。它配备了 block-4 扩散草稿器,并提供了完整的训练代码以及训练/开发/测试数据集。
Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench’s public tiers (0.935 vs 0.866 for Jev). 在未经过训练的测试数据上表现优于 Kev-9B 和 Jev(0.889 对比 0.822 和 0.857),在 JevBench 的公开层级上也表现更佳(0.935 对比 Jev 的 0.866)。
Supports yes/no (noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API. 通过兼容 Jev 的 API,支持在同一请求中处理“是/否”(noul)、多项选择(choice)和评分(score)问题。
About 0.3 s per request without thinking and a 3.3 s median with it on one H100. Can be sped up by truncating chain length. Runs on CUDA (Hopper for the FP8 kernel). 在单张 H100 上,不开启思考时每个请求耗时约 0.3 秒,开启思考时中位数耗时 3.3 秒。可以通过截断推理链长度来加速。运行于 CUDA 环境(FP8 内核需 Hopper 架构)。
Problem
问题
Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides. This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public). 类 Jev 模型虽然能提供校准后的决策概率,但准确率较低。因此,许多流水线依赖推理模型作为后备方案。Jeeves 使用 CISPO 训练了一个类 Jev 的 Qwen3.5-9B 模型(LoRA 和指针头),使其在决策前进行推理。这使得模型在域外任务上表现更好,并在 JevBench Hard(公开)测试中超越了 Jev。
Results
结果
Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes. 开启思考、贪婪搜索、2,560 token 上限下的准确率。Kev-9B 和 Jev 列的数据为 Kev 官方发布。
| bench | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out, item-weighted) | 0.822 | 0.857 | 0.889 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.800 | 0.746 |
| JevBench overall (231 public items) | 0.715* | 0.866 | 0.935 |
| QNLI | 0.925 | 0.925 | 0.913 |
| SciQ | 0.963 | 0.988 | 0.991 |
| TweetEval offensive | 0.775 | 0.813 | 0.813 |
| PAWS | 0.763 | 0.788 | 0.875 |
| MMLU | 0.738 | 0.900 | 0.793 |
| Emotion | 0.600 | 0.588 | 0.647 |
| Held-out rule structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
| MMLU-Pro (10-way) | 0.515 | 0.840 | 0.739 |
| Buried state | 0.740 | 0.700 | 0.759 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0.000 | 0.090 | 0.055 |
| JevBench hard (111 public items) | 0.451* | 0.730 | 0.865 |
| JevBench ECE (public items) | 0.049 | 0.037 | - |
* No Kev-9B JevBench result is published. These are Kev-8B (Qwen3). All JevBench numbers are on the public easy, standard and hard tiers (231 items). The sealed judge tier is not included, and the Jev and Kev numbers are restricted to the same public items. Without thinking the same checkpoint scores 0.804 on our test split (2,962 items), against 0.840 with it.
* 未发布 Kev-9B 的 JevBench 结果。此处为 Kev-8B (Qwen3) 的数据。所有 JevBench 数据均基于公开的简单、标准和困难层级(231 个项目)。不包含密封评判层级,且 Jev 和 Kev 的数据仅限于相同的公开项目。在我们的测试集(2,962 个项目)上,同一检查点在不思考时得分为 0.804,开启思考时为 0.840。
Quickstart
快速开始
Requirements: Python 3.12 and a CUDA GPU. 要求:Python 3.12 和 CUDA GPU。
pip install -r requirements.txt
Download the released weights and serve them: 下载已发布的权重并进行部署:
hf download PostHog/jeeves --local-dir jeeves-weights
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009
Or fuse your own trained checkpoint into a standalone model and serve it with a drafter: 或者将你自己训练的检查点融合为独立模型,并配合草稿器部署:
python export.py runs/cispo/final --out runs/fused
python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009
Then send a request in Jev’s format: 然后以 Jev 的格式发送请求:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"returns": "Exchanges, refunds, wrong or damaged items", "shipping": "Delivery status, delays, lost packages", "billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?", "criteria": ["Calm", "Frustrated", "Very angry"]}
},
"options": {"max_think": 512}}'
Python SDK
Python SDK
sdk/ is a drop-in replacement for Jev’s Python SDK (typesafe-sdk):
sdk/ 是 Jev Python SDK (typesafe-sdk) 的直接替代品:
pip install ./sdk
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
result = client.system_one(
state="I was charged twice. Please help.",
questions={
"billing": Noul(instructions="Is this about billing?"),
"tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
"urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
},
max_think=768,
return_reasoning=True,
)
print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)
The client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s. 客户端默认连接至 http://127.0.0.1:8009(或 JEEVES_BASE_URL),无需 API 密钥,等待时间最长为 120 秒。
Options
选项
options is optional and ignored by Jev clients that don’t send it. Server-wide defaults are set with the matching serve flags.
options 是可选的,不发送该参数的 Jev 客户端会忽略它。服务器范围的默认值通过相应的部署标志设置。
| option | default | effect |
|---|---|---|
| think | true | false answers from the prompt alone (about 0.3 s) |
| max_think | 2560 | truncates each reasoning chain at this many tokens, then answers |
| nothink_threshold | null | answers without thinking when the no-think confidence is at least this value |
| return_reasoning | false | adds each question’s reasoning text to the response |
| 选项 | 默认值 | 效果 |
|---|---|---|
| think | true | false 表示仅根据提示词回答(约 0.3 秒) |
| max_think | 2560 | 将每个推理链截断为此 token 数,然后回答 |
| nothink_threshold | null | 当不思考的置信度至少达到此值时,直接回答而不进行思考 |
| return_reasoning | false | 将每个问题的推理文本添加到响应中 |
How it works
工作原理
Questions, states and answers are loaded into the Qwen chat template like <state> …state… <q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <think>. The model then rolls out its reasoning chain, and after the </think> token we append </think> <q> instructions <opt> option.
问题、状态和答案被加载到 Qwen 聊天模板中,格式如下:<state> …state… <q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <think>。随后模型展开其推理链,在 </think> token 之后,我们追加 </think> <q> instructions <opt> option。