Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

面向资源受限边缘设备的可靠推理:神经符号路由技术

Abstract: Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems.

摘要: 在边缘硬件上运行语言模型可以在无需网络连接的情况下提供私密且低延迟的推理能力。然而,能够适配此类设备的小型模型在处理计算机本应擅长的任务(如算术、代数和形式逻辑问题)时,往往表现得不可靠。

We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit.

我们认为,这种不可靠性在很大程度上是可以避免的。许多看似需要推理的查询实际上在结构上是确定性的,并允许快速且精确的符号化求解。因此,强迫概率模型去近似这些问题,不仅牺牲了准确性和能源,且收效甚微。

We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems.

我们提出了一种神经符号路由器,它能够对每个传入的查询进行分类,并将其分发给成本最低的正确求解器:将结构化任务发送给确定性引擎,而将小型语言模型(SLM)保留用于处理开放式的应用题。

Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle.

我们没有采用手动编写路由逻辑的方式,而是利用 L* 语法推理算法学习了一个确定性有限自动机(DFA),其中以 SLM 作为成员资格预言机(membership oracle),以标注数据作为等价性预言机(equivalence oracle)。

On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers.

在树莓派 4B(8GB 内存,无 GPU)上,通过对来自 DeepMind Mathematics、GSM8K 和 RuleTaker 的 100 个未测试提示词进行评估,学习到的路由方案实现了 100% 的路由准确率和 98.3% 的整体准确率(在 512 token 的推理预算下,应用题准确率为 93.3%)。相比之下,最强的基准代理 Program-of-Thought 的准确率为 72.0%,而使用相同求解器的工具调用代理准确率为 58.7%。

Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.

由于格式化查询无需经过模型处理,路由器可在 1-11 毫秒内完成响应;在 30-token 的配置下,其运行速度比 Program-of-Thought 快 8.8 倍,能效提升了 2.8 倍。