AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

AnchorSIPS:用于证据支持的精神病风险症状测量的合成数据集与评估资源

Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets.

摘要: 人工智能在精神病风险评估方面的进展受到数据获取瓶颈的限制。由于隐私、管理和知情同意等方面的限制,真实的临床访谈记录难以共享。我们提出了 AnchorSIPS,这是一个包含 1 万份结构化精神病风险访谈的合成数据集,并配有基于转录文本的测量目标。

Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms (“frank psychosis”), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms.

每份访谈均以 Mini-SIPS(一种由临床医生执行的精神病风险访谈)为模型。它涵盖了病史、24 个症状问题、患者确认项目的后续证据、关于类妄想症状(异常信念)、类幻觉症状(异常感知)和沟通紊乱的判断、对明确精神病水平症状(“显性精神病”)的排除,以及最终的减弱型精神病综合征(APS)诊断——这是一种表现为较轻或早期精神病症状的高风险状态。

The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient’s clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair.

APS 诊断并非一个独立的标签。它取决于早期的确认情况、支持性的后续细节、症状类别的判断以及对显性精神病的排查。每一个中间决策都锚定在其支持性的转录对话轮次上。AnchorSIPS 是通过“先规划,后实现”的流程生成的。隐藏的病例表规定了患者的临床状态,确定性规划器固定了访谈结构,而大语言模型(LLM)仅在验证和有限修复的前提下实现患者的言语表达。

Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.

在生成前固定标签和结构,避免了多轮 LLM 对话中常见的轮次间不一致问题。在七个 LLM 基准测试中,模型能够恢复粗略的决策,但在提取后续细节或引用支持性转录轮次方面表现不佳,因此最终标签的性能表现往往高估了访谈的实际能力。AnchorSIPS 旨在用于证据提取、基于转录的测量以及部分披露情况下的不确定性研究。