Text2Dashboard: A Governed Agent Architecture for Natural-Language Dashboard Generation over Enterprise DataBrain
Computer Science > Artificial Intelligence arXiv:2610.06914 (cs) [Submitted on 2 Oct 2026] Title: Text2Dashboard: A Governed Agent Architecture for Natural-Language Dashboard Generation over Enterprise DataBrain Authors: Yiou Wu, Zezhi Tang, Ningwei Bai, Liuhaichen Yang.
计算机科学 > 人工智能 arXiv:2610.06914 (cs) [提交于 2026 年 10 月 2 日] 标题:Text2Dashboard:一种用于企业 DataBrain 自然语言仪表板生成的受控代理架构 作者:Yiou Wu, Zezhi Tang, Ningwei Bai, Liuhaichen Yang。
Abstract: Text2Dashboard is a DataBrain-specific prototype that turns natural-language analytic requests into inspectable dashboards. An installable Codex plugin and standalone Agent Runtime combine schema-constrained model decisions with typed tools, persistent state, and deterministic Hooks for approval, audit, checkpointing, recovery, and failure handling.
摘要:Text2Dashboard 是一个针对 DataBrain 的原型系统,它能将自然语言分析请求转换为可检查的仪表板。一个可安装的 Codex 插件和独立的代理运行时(Agent Runtime)将受模式约束的模型决策与类型化工具、持久化状态以及用于审批、审计、检查点、恢复和故障处理的确定性钩子(Hooks)相结合。
The pipeline resolves entities, discovers metadata, enforces read-only SQL, composes dashboards, and applies static checks, dynamic preflight, and browser inspection. The model proposes actions while deterministic software controls execution and records state transitions.
该流水线负责解析实体、发现元数据、强制执行只读 SQL、构建仪表板,并应用静态检查、动态预检和浏览器检查。模型负责提出操作建议,而确定性软件则控制执行过程并记录状态转换。
We evaluate the workflow on frozen real-DataBrain tasks and controlled Hook faults. Strict success was 6/8 on metadata and SQL tasks: metadata selection passed 4/4, all four SQL tasks met semantic criteria, and 2/4 met the exact output-column contract.
我们在冻结的真实 DataBrain 任务和受控的钩子故障上评估了该工作流。在元数据和 SQL 任务中,严格成功率为 6/8:元数据选择通过了 4/4,所有四个 SQL 任务均符合语义标准,其中 2/4 符合精确的输出列契约。
The final release passed 4/4 single-panel dashboard tasks, one two-panel task, and one existing-dashboard refinement; a parameterised task exceeded its step limit. All ten fault scenarios met their specified outcomes without unapproved external side effects.
最终版本通过了 4/4 的单面板仪表板任务、一个双面板任务以及一个现有仪表板的优化任务;一个参数化任务超出了其步数限制。所有十种故障场景均达到了预期的结果,且没有产生未经批准的外部副作用。
Model inference accounted for over 97% of observed runtime in every reported group. These small, DataBrain-specific results do not establish production readiness, general text-to-SQL accuracy, or an efficiency advantage over manual dashboard construction.
在每个报告组中,模型推理占用了超过 97% 的观测运行时间。这些针对 DataBrain 的小规模结果并不能证明其已具备生产就绪性、通用的文本转 SQL 准确性,或相比手动构建仪表板的效率优势。