VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes

VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes

VERGE:用于临床记录中早发性结直肠癌症状基础提取的验证增强细化方法

Abstract: Early-onset colorectal cancer is increasing among younger adults, yet red-flag symptoms in this age group have no evidence-based guidelines for follow-up testing, and structured encounter data do not capture the detail needed to support early detection and inform follow-up, including symptom duration, context, and family history, an established colorectal-cancer risk factor.

摘要: 早发性结直肠癌在年轻人中的发病率正在上升,但针对该年龄段的“红旗”症状尚无循证指南来指导后续检查。此外,结构化的就诊数据无法捕捉支持早期发现和后续随访所需的详细信息,例如症状持续时间、背景以及作为结直肠癌既定风险因素的家族史。

This study aimed to develop and evaluate an automated method for extracting six red-flag symptoms and family-history risk status from free-text clinical notes. We developed VERGE, an agentic workflow in which an initial label and evidence are proposed using retrieval-augmented generation, then passed through a bounded verification-refinement cycle that checks textual grounding and clinical validity, corrects and rechecks a claim until resolved or a limit is reached, and escalates unresolved claims for human review.

本研究旨在开发并评估一种自动化方法,用于从自由文本临床记录中提取六种“红旗”症状及家族史风险状态。我们开发了 VERGE,这是一个智能体工作流:首先利用检索增强生成(RAG)提出初步标签和证据,随后通过一个有界的“验证-细化”循环进行处理,该循环会检查文本依据和临床有效性,对声明进行修正和复核,直至问题解决或达到限制次数,并将无法解决的声明升级至人工审核。

VERGE was evaluated on 4,033 clinician-labeled note-finding pairs against a single-agent baseline, a rule-based clinical language-processing baseline, and an alternative underlying language model. Compared with the single-agent baseline, VERGE reduced false positive findings, improving precision from 0.764 to 0.849 and MCC from 0.681 to 0.730, a balanced gain across the precision-recall trade-off, and resolved most flagged errors autonomously, with human review required for only 1.5 percent of claims.

VERGE 在 4,033 对由临床医生标注的记录-发现对上进行了评估,并与单智能体基线、基于规则的临床语言处理基线以及另一种底层语言模型进行了对比。与单智能体基线相比,VERGE 减少了假阳性发现,将精确率从 0.764 提高到 0.849,MCC 从 0.681 提高到 0.730,在精确率与召回率之间取得了平衡的提升,并自主解决了大部分标记错误,仅有 1.5% 的声明需要人工审核。

These results indicate that a bounded, verification-based workflow can reduce unnecessary positive findings without sacrificing the ability to detect true ones. This approach offers a path toward more reliable and trustworthy clinical language-processing tools to support colorectal cancer risk assessment in younger patients.

这些结果表明,一种有界的、基于验证的工作流可以在不牺牲真实病例检测能力的前提下,减少不必要的阳性发现。该方法为开发更可靠、更值得信赖的临床语言处理工具提供了一条路径,以支持年轻患者的结直肠癌风险评估。