Native vs OCR PDF Text in Node.js: Choose Page Indexing by Ownership

Native vs OCR PDF Text in Node.js: Choose Page Indexing by Ownership

Node.js 中的原生 PDF 文本与 OCR:根据所有权选择页面索引策略

TL;DR: Use born-digital text extraction when your B2B SaaS team owns the PDF templates; preserve the extractor’s page boundary and index only redacted text. Choose OCR when customers control the templates, scans are valid input, or extraction quality cannot be enforced. Record which path produced every page so citations remain explainable. 简而言之:当你的 B2B SaaS 团队拥有 PDF 模板时,请使用原生数字文本提取;保留提取器的页面边界,并仅索引脱敏后的文本。当客户控制模板、扫描件作为有效输入或无法保证提取质量时,请选择 OCR。记录每个页面的生成路径,以便引用内容可追溯。

Input and ownershipPickWhyRequired signal
Your team owns the export templateNative PDF extractionTemplate changes can be testedExtracted characters and empty-page count
Customer uploads may be scansOCRImage-only pages lack a useful text layerOCR path, confidence, and page count
Mixed portfolioNative first, then explicit OCR fallbackKnown-good exports avoid OCR while scans remain validExtractor kind per page and fallback reason
输入与所有权选择原因所需信号
团队拥有导出模板原生 PDF 提取模板变更可测试提取字符数与空页计数
客户上传可能包含扫描件OCR纯图片页面缺乏有效的文本层OCR 路径、置信度与页码
混合文档组合优先原生,后置 OCR 回退优质导出件避免 OCR,同时兼容扫描件每页的提取器类型与回退原因

The firm rule is about control. Template ownership lets a team test the text layer before release. Without that control, visual correctness does not prove extractable text. For a document-sharing workflow, the pipeline is PDF ingestion, page-preserving extraction, personal-data redaction, chunking, embedding, and retrieval. The index never receives unredacted page text. 核心原则在于控制权。模板所有权允许团队在发布前测试文本层。若缺乏这种控制,视觉上的正确性并不能证明文本可提取。对于文档共享工作流,流水线应包括:PDF 摄入、保留页面的提取、个人数据脱敏、分块、向量化嵌入和检索。索引库绝不应接收未经脱敏的页面文本。

How should Node.js index PDF text per page? The extraction adapter should. A PDF is a structured document defined by ISO 32000-2, not a bag of newline-delimited strings. If an application concatenates the whole file and guesses page breaks later, a retrieval hit cannot reliably point back to the page a reviewer must inspect. Treat pageNumber as provenance. Keep it 1-based at the adapter boundary, validate it, and carry it through redaction, chunking, storage, and result rendering. One convention removes an ugly class of off-by-one citations. Node.js 应如何按页面索引 PDF 文本?应由提取适配器完成。PDF 是由 ISO 32000-2 定义的结构化文档,而非一堆以换行符分隔的字符串。如果应用程序将整个文件拼接在一起,事后再猜测分页位置,检索结果将无法可靠地指向审查员需要检查的页面。将 pageNumber 视为数据来源(provenance)。在适配器边界处保持从 1 开始计数,进行验证,并将其贯穿于脱敏、分块、存储和结果渲染的全过程。这一约定消除了“差一错误”(off-by-one)引用带来的麻烦。

Page numbers matter. Text may appear in an order different from the order a person reads it. Multi-column layouts, headers, footers, and positioned glyphs expose that gap. Template owners can prevent regressions with fixture PDFs and expected page text. For customer-controlled files, use quality gates and route weak pages to OCR or review rather than pretending every successful parse is useful. 页码至关重要。文本出现的顺序可能与人类阅读顺序不同。多栏布局、页眉、页脚和定位字形都会暴露这种差异。模板所有者可以通过测试用例 PDF 和预期的页面文本来防止回归。对于客户控制的文件,请使用质量门禁,并将质量较差的页面路由至 OCR 或人工审核,而不是假装每一次成功的解析都是有效的。

Diagram in words: upload enters quarantine; a parser emits numbered pages; a redactor replaces personal fields; a chunker emits page-scoped passages; the retrieval index stores those passages plus provenance; the sharing service renders only redacted results. Raw text takes a separate, access-controlled path to deletion. 流程图解:上传文件进入隔离区;解析器输出带页码的页面;脱敏器替换个人字段;分块器输出页面范围内的段落;检索索引存储这些段落及来源信息;共享服务仅渲染脱敏后的结果。原始文本则通过独立的、受访问控制的路径进行删除。

Choose this path for invoices, account summaries, or reports generated from templates your team releases. The advantage is testability. Keep representative PDFs as fixtures, assert the page count, and assert stable anchor phrases on each page after every template change. A test should fail if page 3 becomes empty even when the rendered PDF still looks fine. Count pages with no extracted characters. Track the ratio of redacted characters to extracted characters as a distribution, without logging the characters themselves. Alert on a sustained change by template version, because a global average can hide one broken export family. 对于由团队模板生成的发票、账户摘要或报告,请选择此路径。其优势在于可测试性。保留代表性 PDF 作为测试用例,断言页码数量,并在每次模板变更后断言每页的稳定锚点短语。如果第 3 页变为空白,即使渲染出的 PDF 看起来正常,测试也应失败。统计没有提取出字符的页面。将脱敏字符与提取字符的比例作为分布进行跟踪,但不要记录字符本身。针对模板版本的持续变化发出警报,因为全局平均值可能会掩盖某个损坏的导出系列。

Never log source text. OWASP’s logging guidance calls out data that commonly needs removal, masking, sanitization, hashing, or encryption, including personal data. Here, document identifiers should be pseudonymous, and trace data should describe stages and counts rather than document contents. 切勿记录源文本。OWASP 的日志记录指南指出了通常需要删除、掩码、清洗、哈希或加密的数据,包括个人数据。在此场景下,文档标识符应使用化名,追踪数据应描述处理阶段和计数,而非文档内容。

Pick OCR when uploads vary. Choose OCR when valid inputs include scanned pages or another organization owns document creation. It adds an error surface, needs its own quality signal, and may produce plausible text that is wrong. A low-confidence page should not produce a confident-looking citation. Keep the fallback visible. Store ocr as the extractor kind, retain page-level confidence when the engine supplies it, and define a review state below your acceptance threshold. Derive that threshold from labeled documents in the languages and layouts you accept. Do not borrow a magic number from a demo. Scans change that. 当上传内容多变时选择 OCR。当有效输入包含扫描件或文档由其他组织创建时,请选择 OCR。它增加了错误面,需要独立的质量信号,且可能产生看似合理但错误的文本。低置信度的页面不应产生看起来很确定的引用。保持回退机制的可见性。将 ocr 存储为提取器类型,在引擎提供时保留页面级置信度,并定义低于验收阈值的审核状态。该阈值应根据你所接受的语言和布局的标注文档得出,不要直接照搬演示代码中的“魔术数字”。扫描件会改变这一切。

For mixed PDFs, decide page by page. A ten-page upload can contain nine born-digital pages and one scanned signature page. Preserve all ten page numbers even if one yields no searchable chunk after redaction. Gaps are evidence. Consider a concrete upload with pages 1 through 10. Pages 1 through 6 contain selectable account text, page 7 is a scanned authorization form, and pages 8 through 10 return to selectable text. Native extraction is the right first path, but it is not suitable for page 7. OCR is the right fallback there, but running OCR over all ten pages would replace known text with probabilistic output for no retrieval benefit. The trade-off is deliberate: a page-level branch makes the job state and monitoring more complex, while preserving the strongest available evidence for each citation. The index should still contain one provenance convention across both branches. A result from page 7 carries extractor: "ocr"; a result from page 8 carries extractor: "native". Reviewers can now distinguish those evidence paths without seeing private source text in logs. 对于混合型 PDF,请逐页决策。一份十页的上传文件可能包含九页原生数字页面和一页扫描的签名页。即使其中一页在脱敏后没有可搜索的块,也要保留所有十个页码。缺失本身就是一种证据。考虑一份 1 到 10 页的上传文件:1 到 6 页包含可选取的账户文本,第 7 页是扫描的授权表,8 到 10 页恢复为可选取的文本。原生提取是首选路径,但不适用于第 7 页。OCR 是该页的正确回退方案,但对全部十页运行 OCR 会用概率性输出替换已知文本,且对检索毫无益处。这种权衡是刻意的:页面级分支使作业状态和监控变得复杂,但为每次引用保留了最强的可用证据。索引库应在两个分支中保持统一的来源约定。来自第 7 页的结果标记为 extractor: "ocr";来自第 8 页的结果标记为 extractor: "native"。审查员现在无需在日志中查看私有源文本,即可区分这些证据路径。

Implement a redaction-first page contract. The parser or OCR adapter returns numbered pages; the policy layer owns redaction; the indexer accepts only redacted chunks. This TypeScript example leaves parsing, OCR, and embedding behind interfaces, so the contract does not depend on one product. 实现“脱敏优先”的页面契约。解析器或 OCR 适配器返回带页码的页面;策略层负责脱敏;索引器仅接收脱敏后的块。以下 TypeScript 示例将解析、OCR 和嵌入封装在接口之后,因此契约不依赖于特定的产品。

type ExtractorKind = "native" | "ocr";

type ExtractedPage = {
  pageNumber: number;
  text: string;
  extractor: ExtractorKind;
  confidence?: number;
};

type IndexedChunk = {
  documentId: string;
  pageNumber: number;
  chunkNumber: number;
  text: string;
  extractor: ExtractorKind;
};

interface PageExtractor {
  extract(pdf: Uint8Array): Promise<ExtractedPage[]>;
}

interface RetrievalIndex {
  upsert(chunks: IndexedChunk[]): Promise<void>;
}

type Redact = (text: string) => string;

function chunkPage(text: string, maxChars = 1_200): string[] {
  const paragraphs = text.split(/\n\s*\n/).map((part) => part.trim()).filter(Boolean);
  const chunks: string[] = // ...