Multimodal open d1 decision models for the edge
本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。
Back to Articles Multimodal open d1 decision models for the edge Team Article Published October 7, 2026 Upvote 17 +11 Aurelien Lac Aurelien-Lac Follow LiquidAI Fernando Fernandes Neto fernandofernandes Follow LiquidAI Edoardo Mosca EdoardoMosca Follow LiquidAI Maxime Labonne mlabonne Follow LiquidAI Leonie Monigatti iamleonie Follow LiquidAI
返回文章:面向边缘计算的多模态开源 d1 决策模型。团队文章,发布于 2026 年 10 月 7 日,点赞 17 次,+11。作者:Aurelien Lac、Fernando Fernandes Neto、Edoardo Mosca、Maxime Labonne、Leonie Monigatti(均来自 LiquidAI)。
Today, we release two open decision models in our d1 decision model family: d1-3B and d1-omni-600M (experimental). Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11). Multimodal: d1-3B supports text and images, while d1-omni-600M supports text and images or text and audio. Fast: d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano.
今天,我们发布了 d1 决策模型家族中的两款开源决策模型:d1-3B 和 d1-omni-600M(实验性)。在 Decision Index 0.2.1 中,d1-3B 是 10B 参数以下表现最好的决策模型,得分 48.57,领先于所有 4B 和 9B 模型以及 Decider 35B-A3B(47.11)。多模态方面:d1-3B 支持文本和图像,而 d1-omni-600M 支持文本与图像或文本与音频。速度方面:d1-3B 在 NVIDIA Jetson AGX Thor 上回答一个问题仅需 16 毫秒,在 Jetson AGX Orin 上为 26 毫秒,在 Jetson Orin Nano 上为 50 毫秒。
How we built decision models for the edge: These open d1 decision models are built on our Liquid Foundation Models (LFMs). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass. d1-3B and d1-omni-600M are trained from two very different backbones: d1-3B is trained from LFM2.5-VL-3B, our latest VLM, which is decoder-only. It accepts text and images as inputs. d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder. It adds vision and audio encoders to handle all three modalities. It accepts either text and image, or text and audio as inputs. This model is currently in an early research release and is undergoing further development.
我们如何构建边缘决策模型:这些开源 d1 决策模型基于我们的 Liquid 基础模型 (LFM)。与我们的生成式模型不同,决策模型不生成 token,而是在单次前向传递中给出答案。d1-3B 和 d1-omni-600M 基于两种截然不同的主干网络训练:d1-3B 基于我们最新的仅解码器架构 VLM——LFM2.5-VL-3B 训练,接受文本和图像作为输入。d1-omni-600M 基于双向编码器 LFM2.5-Encoder-350M 训练,增加了视觉和音频编码器以处理三种模态,接受文本与图像或文本与音频作为输入。该模型目前处于早期研究发布阶段,正在进一步开发中。
Benchmark results: We benchmarked d1-3B and d1-omni-600M on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. d1-3B achieves a mean score of 82.9, the highest in the table and above Decider 4B. d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.
基准测试结果:我们在涵盖阅读理解、毒性检测、意图分类、医学问答和跨语言理解的七个公共数据集上对 d1-3B 和 d1-omni-600M 进行了基准测试。d1-3B 的平均得分为 82.9,是表中的最高分,超过了 Decider 4B。d1-omni-600M 得分为 78.4,仅用四分之一的参数就超过了 Decider 2B (77.1)。
We validated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, and that d1-omni-600M handles all three modalities. We do not report any vision or audio benchmarks, as the Decision Index v0.3 includes only a private vision split and audio decision benchmarks are currently an open problem.
我们验证了 d1-3B 在标准视觉基准测试中保留了其 LFM2.5-VL-3B 主干的视觉能力,并且 d1-omni-600M 能够处理所有三种模态。我们没有报告任何视觉或音频基准测试结果,因为 Decision Index v0.3 仅包含私有视觉拆分,且音频决策基准测试目前仍是一个未解决的问题。
Speed: In collaboration with NVIDIA, we evaluated d1-3B on the NVIDIA stack across NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Since d1-omni-600M is an early research release, we don’t report any speed numbers for it in this release. Edge inference: d1-3B answers a single question in under 50 ms on every measured device. Three questions take only 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms.
速度:我们与 NVIDIA 合作,在 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 和 Jetson Orin Nano 等 NVIDIA 平台上评估了 d1-3B。由于 d1-omni-600M 属于早期研究发布,本次不报告其速度数据。边缘推理:d1-3B 在所有测试设备上回答单个问题的时间均低于 50 毫秒。处理三个问题的时间仅为处理一个问题的 1.3 倍,其中 AGX Thor 从 16 毫秒增加到 20 毫秒。
GPU inference: On GPU, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms.
GPU 推理:在 GPU 上,d1-3B 在两个平台上回答一个问题的时间均低于 10 毫秒,处理 384px 图像的时间均低于 18 毫秒。
How to use open d1 decision models: Reach for d1 decision models when you need fast, structured decisions, including multimodal inputs. d1-3B delivers the highest decision quality at its size, while d1-omni-600M fits where footprint matters. Install the dependencies (requires transformers>=5.14): pip install “transformers>=5.14” torch torchvision pillow. These model ship their own code, so load it with trust_remote_code=True.
如何使用开源 d1 决策模型:当您需要快速、结构化的决策(包括多模态输入)时,请选择 d1 决策模型。d1-3B 在同等规模下提供最高的决策质量,而 d1-omni-600M 则适用于对资源占用敏感的场景。安装依赖项(需要 transformers>=5.14):pip install “transformers>=5.14” torch torchvision pillow。这些模型自带代码,因此请使用 trust_remote_code=True 进行加载。
Get Started with open d1 decision models: Both decision models are open-weight and available on Hugging Face today: Download: d1-3B and d1-omni-600M on Hugging Face. Try: run the demos in our System One Arcade Hugging Face Space. We can’t wait to see what you build. Citation: If you use this work, please cite the release blog: @a
开始使用开源 d1 决策模型:这两款决策模型均为开源权重,现已在 Hugging Face 上线:下载:在 Hugging Face 上获取 d1-3B 和 d1-omni-600M。尝试:在我们的 System One Arcade Hugging Face Space 中运行演示。我们迫不及待地想看到您的作品。引用:如果您使用此项工作,请引用发布博客:@a