Invoice Rule Evidence: 5 PDF Model Field Extraction Decisions
Invoice Rule Evidence: 5 PDF Model Field Extraction Decisions
发票规则证据:PDF 模型字段提取的 5 个决策
TL;DR: For marketplace invoice ingestion, choose rules when the document family is controlled and signatures must be validated against the original bytes. Choose model-based extraction when layouts and labels vary, but keep it away from signature validation and require field-level evidence. In a mixed marketplace, the practical design is usually a cascade: preserve the source PDF, verify its signature, attempt deterministic extraction, route uncertain fields to a model, and record every decision in an append-only audit event.
简而言之: 对于市场发票接入,当文档格式受控且必须根据原始字节验证签名时,请选择规则提取。当布局和标签多变时,请选择基于模型的提取,但不要将其用于签名验证,并要求提供字段级的证据。在混合型市场中,实用的设计通常是级联式:保留原始 PDF,验证其签名,尝试确定性提取,将不确定的字段路由至模型,并将每个决策记录在仅追加(append-only)的审计事件中。
| Pick this path | Best fit | Main failure mode | Evidence to retain |
|---|---|---|---|
| 选择此路径 | 最适用场景 | 主要故障模式 | 需保留的证据 |
| Rules | A few stable seller templates | Silent layout drift | Rule version, page, bounding box, raw value |
| 规则 | 少量稳定的卖家模板 | 静默布局偏移 | 规则版本、页码、边界框、原始值 |
| Model | Many unknown layouts or scans | Plausible but unsupported values | Model version, prompt/schema version, page image, confidence |
| 模型 | 大量未知布局或扫描件 | 看似合理但无支撑的值 | 模型版本、提示词/模式版本、页面图像、置信度 |
| Cascade | Mixed marketplace traffic | Bad routing or thresholds | Every stage result and routing reason |
| 级联 | 混合市场流量 | 路由或阈值错误 | 每个阶段的结果和路由原因 |
Accuracy alone is the wrong finish line. An invoice total that looks right but cannot be traced to page 2, or one extracted from a PDF whose signed byte range was altered, is weak evidence. Fidelity means preserving what the document said, where it said it, and which transformation produced the normalized value.
仅追求准确率是错误的终点。一个看起来正确但无法追溯到第 2 页的发票总额,或者从签名字节范围已被篡改的 PDF 中提取出的数据,都是薄弱的证据。保真度意味着要保留文档说了什么、在何处说的,以及通过何种转换产生了标准化后的数值。
1. Should rule-based PDF parsing or model field extraction run first?
1. 应该先运行基于规则的 PDF 解析还是模型字段提取?
Start with document diversity, not fashion. A rule-based parser can be exact when sellers emit PDFs from a stable template: locate the text token near “Invoice total”, constrain its region, parse the currency, and verify the arithmetic. Its behavior is inspectable. A changed font, shifted column, flattened form, or scanned page can invalidate those assumptions without making the file invalid.
从文档的多样性出发,而不是盲目跟风。当卖家使用稳定的模板生成 PDF 时,基于规则的解析器可以做到精确:定位“Invoice total”附近的文本标记,约束其区域,解析货币,并验证算术结果。其行为是可审查的。字体更改、列偏移、扁平化表单或扫描页面可能会使这些假设失效,但不会使文件本身无效。
A model-based extractor handles label and layout variation better because it can map semantically similar phrases into one schema. It also introduces a different risk: a syntactically valid answer can lack document support. Never let schema validation masquerade as evidence validation. For a marketplace, make the first attempt with rules only when a template fingerprint is known and the required text layer is present. Route everything else to model extraction. This keeps deterministic work deterministic while giving unfamiliar invoices a path forward.
基于模型的提取器能更好地处理标签和布局的变化,因为它能将语义相似的短语映射到同一个模式中。它也带来了另一种风险:语法上正确的答案可能缺乏文档支持。永远不要让模式验证伪装成证据验证。对于市场平台,仅在已知模板指纹且存在所需文本层时,才优先尝试规则提取。将其余所有内容路由至模型提取。这既保持了确定性工作的确定性,又为不熟悉的格式提供了处理路径。
2. Preserve signature truth before touching content
2. 在处理内容前保留签名的真实性
PDF digital signatures cover byte ranges in the original file. Parsing, rendering, optimizing, or regenerating the PDF before verification can destroy the exact artifact needed to assess that signature. So the source bytes come first. Store a cryptographic digest, verification result, signer certificate details, and verification time alongside the immutable object identifier. This separation matters. Signature verification answers whether covered bytes and credentials validate under the chosen policy. Field extraction answers what those bytes appear to contain. Neither answer proves the other.
PDF 数字签名覆盖了原始文件中的字节范围。在验证前对 PDF 进行解析、渲染、优化或重新生成,可能会破坏评估该签名所需的精确工件。因此,原始字节必须优先处理。将加密摘要、验证结果、签名者证书详情和验证时间与不可变的对象标识符一起存储。这种分离至关重要。签名验证回答的是所覆盖的字节和凭据在选定策略下是否有效;字段提取回答的是这些字节看起来包含什么。两者互不证明。
3. Make every field carry its own receipt
3. 让每个字段都携带自己的“收据”
One document-level confidence score hides the exact disagreement an operator needs to inspect. Represent each extracted field as a claim with provenance. Do not overwrite sourceText with the normalized value. If the invoice prints “1.234,50 EUR”, the audit record needs that literal text even if downstream accounting receives 1234.50 and EUR. Page numbers and bounding boxes let a reviewer return to the visual source instead of trusting a parser log.
单一的文档级置信度分数会掩盖操作员需要检查的具体分歧。将每个提取的字段表示为带有来源证明的声明。不要用标准化后的值覆盖 sourceText。如果发票上打印的是“1.234,50 EUR”,即使下游会计系统接收的是 1234.50 和 EUR,审计记录也必须保留该字面文本。页码和边界框让审查者能够回到视觉源文件,而不是盲目信任解析器日志。
4. Test disagreement, drift, and abstention
4. 测试分歧、偏移和弃权
Build the test set around failure classes. Include native-text PDFs, scans, rotated pages, duplicate labels, credit notes, multi-currency orders, line items split across pages, and signed revisions. Keep expected values and expected evidence locations. A correct value from the wrong label should fail. Measure field-level exact match after explicit normalization, but also measure unsupported-claim rate, abstention rate, signature-verification coverage.
围绕故障类别构建测试集。包括原生文本 PDF、扫描件、旋转页面、重复标签、贷记单、多币种订单、跨页拆分的行项目以及已签名的修订版。保留预期值和预期的证据位置。从错误标签中提取出的正确值应被视为失败。在进行显式标准化后测量字段级的精确匹配,同时也要测量无支撑声明率、弃权率和签名验证覆盖率。