FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

FinDialogLens:金融聊天室中用于识别漏单的多方对话事件抽取系统

Abstract: Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants.

摘要: 多方金融聊天室对于销售和交易专业人士至关重要,但其复杂性使得人工恢复漏单变得不可行:每一条报价请求(RFQ)都是一个事件,其最终价格和交易结果往往出现在触发 RFQ 的消息(询价消息)之后的多条消息中,且与来自其他参与者的并发 RFQ 交织在一起。

We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles.

我们将此问题定义为多方对话中的事件抽取(EE)任务,并提出了 FinDialogLens,这是一个混合大模型(LLM)流水线。在该流水线中,精简的微调分类器充当推理时的支架:它们负责检测 RFQ 触发点及价格/交易结果元数据,RFQ 级模块负责分割每个事件的 RFQ 窗口,而交易引擎则负责填充参数角色。

With GPT-4o, FinDialogLens reaches 92.1% and 94.3% accuracy on final price and trade outcome, respectively, outperforming full-chatroom CoT prompting methods; fine-tuned open-source LLMs with as few as 3B parameters achieve comparable performance with modest in-domain data.

在使用 GPT-4o 的情况下,FinDialogLens 在最终价格和交易结果上的准确率分别达到了 92.1% 和 94.3%,优于全聊天室思维链(CoT)提示方法;参数量仅为 3B 的微调开源大模型在少量领域内数据下也能达到相当的性能。

To make the LLM-based solution practical at scale, a difficulty-aware router balances cost and accuracy by allocating RFQs between a low-cost rule-based engine and the higher-performing LLM-powered Trade Engine, cutting LLM calls by 85% on final price while recovering half of the accuracy gap to FinDialogLens (GPT-4o), saving over $300/day at our 70,000-RFQ/day scale.

为了使基于大模型的解决方案在大规模场景下具备实用性,系统引入了一个难度感知路由,通过在低成本的规则引擎和高性能的大模型交易引擎之间分配 RFQ 来平衡成本与准确性。该方法在最终价格识别任务中减少了 85% 的大模型调用量,同时弥补了 FinDialogLens (GPT-4o) 一半的准确率差距,在我们日均 70,000 条 RFQ 的规模下,每天可节省超过 300 美元的成本。