Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Jeff Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code, with the same request format as Jev. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX). Jeff 是基于 Qwen3.5 和 Gemma 4 微调的零样本(zero-shot)分类模型:它们是小巧、快速的决策模型,可直接嵌入你的代码中,并使用与 Jev 相同的请求格式。你只需用简单的语言描述情况并列出选项,Jeff 就能通过单次前向传播为每个选项返回校准后的概率。无需生成文本,无需解析:在 RTX PRO 6000 上单次决策耗时约 22 毫秒,在 Apple M4 Max (MLX) 上约为 28 毫秒。
Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don’t need to appear in the training data; you describe them, and Jeff picks. What it is, and what it isn’t. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won’t match Jev’s, which runs on a much larger model. “零样本”意味着选项可以是任何内容:支持队列、用户意图、审核标签、语音指令或游戏动作。你的分类类别无需出现在训练数据中;你只需描述它们,Jeff 就能进行选择。这就是它的定位。这些模型非常小,能在选项之间做出极快且校准良好的判断,并能轻松集成到本地代码中。在基准测试中,它们的表现接近甚至有时超过 Jev;但受限于体积,其推理能力无法与运行在更大模型上的 Jev 相提并论。
If zero-shot accuracy isn’t good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data. 如果零样本准确率无法满足你的需求,针对自有数据进行简短的微调效果会更好:我们的语音导航微调实验在单张 GPU 上不到半小时,就将留存测试集的准确率从 31.7% 提升到了 95.8%。该项目完全在本地硬件上构建。训练过程在单张 RTX PRO 6000 工作站 GPU 上完成(0.8B 模型训练约 2 小时,2B 模型约 3.5 小时),所有合成训练数据均由开源模型 (Qwen3.8-Flash-Next) 在两台 DGX Sparks 上生成,并在 MacBook 上进行测试。未使用云端 GPU,训练数据中也不包含闭源模型输出;仅使用闭源模型对部分合成数据进行了质量抽检。
Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe. 这是一个独立项目。Jeff 使用与 Jev 相同的请求格式,但它与 Jev 的开发者 TypeSafe 没有任何关联,也未获得其背书。我们的训练代码基于开源的 AutoJev 配方。
Quick start
快速开始
# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
# Apple silicon (MLX, much faster on a Mac; Qwen models only)
uv sync --extra mac
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
Each answer has a probability per option, the chosen option and a confidence. Three question types: choice (pick one of up to 255 options), noul (yes/no, returned as a probability) and score (a point on a scale you describe). Several independent questions in one request are answered together. 每个答案都包含每个选项的概率、选定的选项以及置信度。支持三种问题类型:choice(从最多 255 个选项中选一个)、noul(是/否,以概率形式返回)和 score(在你描述的量表上打分)。一次请求中可以同时回答多个独立问题。
(Note: The original article includes extensive benchmark tables and game testing data. Due to length constraints, the core technical summary is provided above. For the full benchmark data and game performance metrics, please refer to the original source.) (注:原文包含大量基准测试表格和游戏测试数据。受篇幅限制,此处提供核心技术摘要。如需查看完整的基准测试数据和游戏性能指标,请参考原文。)