Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
Vibe Patenting:评估用于专业专利撰写代理的 LLM 裁判
LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. 大语言模型(LLM)裁判正越来越多地被用于评估和改进人工智能生成的输出,但它们在处理复杂专业工作时的可靠性仍不明确。我们通过“Vibe Patenting”研究了这一问题,这是一个用于人工智能代理评估的端到端专利撰写测试平台。
A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate. 一个独立调用的 LLM 裁判会对生成的专利草案进行评估,并提供结构化的反馈以进行迭代修订。在多种发明和撰写代理配置中,由裁判引导的修订始终能提高裁判评估的质量,而未经引导的修订往往会达到瓶颈。
Notably, iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains. 值得注意的是,迭代的裁判反馈使低推理能力的代理能够接近成本高昂得多的高推理能力代理的性能。更强的模型和增强的推理能力通常能提高裁判评估的撰写质量,而特定领域的代理工作流则能带来进一步的提升。
We validate the judge against independent evaluation by a professional patent attorney and find meaningful but strongly metric-dependent agreement and systematic calibration differences. These results highlight both the utility and limitations of LLM judges as evaluators and optimization signals for complex professional workflows. 我们将该裁判的评估结果与专业专利律师的独立评估进行了对比验证,发现两者之间存在有意义但高度依赖指标的一致性,同时也存在系统性的校准差异。这些结果凸显了 LLM 裁判作为复杂专业工作流的评估者和优化信号时的效用与局限性。