Your pull request review time is mostly waiting and rework

Your pull request review time is mostly waiting and rework

你的合并请求(PR)审查时间大多花在等待和返工上

Time to merge is the number teams quote when they say AI review sped things up. It bundles at least three separate costs, and reading the diff is only one of them. When the number drops after adding a tool, the drop usually comes from the other two. “合并耗时”是团队在声称 AI 审查提升效率时最常引用的指标。它至少包含了三个独立的成本,而阅读代码差异(diff)只是其中之一。当引入工具后该数值下降时,通常是因为另外两项成本降低了。

Two sets of measurements make the split concrete. CodeRabbit’s guide to reviewing AI-generated diffs cites LinearB’s 2026 benchmark report: AI-assisted pull requests are 2.6 times larger than unassisted ones, take 4.6 times longer to receive a first review, and have a 30-day acceptance rate of 32.7 percent against 84.4 percent for manual PRs. 两组测量数据让这种拆分变得具体。CodeRabbit 在关于审查 AI 生成代码差异的指南中引用了 LinearB 的 2026 年基准报告:AI 辅助的合并请求(PR)规模是人工 PR 的 2.6 倍,首次审查的等待时间长达 4.6 倍,且 30 天内的接受率仅为 32.7%,而人工 PR 为 84.4%。

CodeRabbit is careful to label those as observed associations rather than causal effects, and the association framing is the honest way to read them. The acceptance gap is the part worth sitting with. If roughly two of every three AI-assisted changes are not accepted within 30 days, the cycle spans several passes rather than one review. CodeRabbit 谨慎地将这些数据标记为“观察到的关联”而非“因果效应”,这种关联性的表述是解读它们的诚实方式。接受率的差距是值得深思的部分。如果大约每三个 AI 辅助的变更中就有两个在 30 天内未被接受,那么整个周期就跨越了多次迭代,而非单次审查。

The second set is narrower and better instrumented. Obada Kraishan’s study of five autonomous coding agents covers 37,623 provenance-labeled PRs across 2,807 repositories, with the pipeline code released for replication. It is a preprint, not peer reviewed. 第二组数据范围更窄且测量更精准。Obada Kraishan 对五个自主编程智能体的研究涵盖了 2,807 个代码库中的 37,623 个带有来源标签的 PR,并发布了用于复现的流水线代码。这是一份预印本,尚未经过同行评审。

Review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, and Claude Code PRs waited a median of 12.6 hours for a first human review. Post-merge outcomes differ by vendor too, with Codex PRs reverted about half as often as human PRs, 6.1 percent against 11.5 percent, while Devin PRs were reverted more often at 14.5 percent. 审查工作量的分布并不均匀:Copilot 的 PR 引发了最多的人工审查和变更请求,而 Claude Code 的 PR 平均需要等待 12.6 小时才能获得首次人工审查。合并后的结果也因供应商而异:Codex PR 的回滚率约为人工 PR 的一半(6.1% 对 11.5%),而 Devin PR 的回滚率则更高,达到 14.5%。

That 12.6-hour median is the number to keep. Reviewers do not read diffs for twelve hours. The wait is queue time: an agent opens a PR, and nothing happens until a human picks it up. A tool that reads faster than a human shortens the few minutes at the end of that window and leaves the twelve hours at the front of it untouched. 那个 12.6 小时的中位数是关键所在。审查者并不会花 12 小时去阅读代码差异。这种等待是“排队时间”:智能体提交了 PR,但在人类介入之前什么都不会发生。一个阅读速度比人类快的工具,只能缩短窗口期末尾的那几分钟,而对前端那 12 小时的等待毫无影响。

Teams that buy review speed by buying reading speed often watch the merge metric barely move, and this is why. The 441 percent increase in code-review time that shows up in the vibe coding literature sits on top of a queue that was already the slow part. 那些试图通过提升阅读速度来换取审查速度的团队,往往会发现合并指标几乎没有变化,原因就在于此。“氛围编程”(vibe coding)文献中提到的代码审查时间增加 441%,实际上是叠加在原本就已经很缓慢的排队环节之上的。

Wait, read, and rework are different numbers. Split the cycle at four timestamps: when the PR opens, when the first review lands, when the first change request lands, and when it is approved and merged. Wait time is the gap between the first pair. Reading time is between the second and third. Rework is everything after a change request, including the author’s fix, the re-review, and any further round trip. 等待、阅读和返工是不同的指标。将周期拆分为四个时间点:PR 开启时、首次审查完成时、首次变更请求发出时,以及最终批准合并时。等待时间是前两个时间点之间的间隔。阅读时间是第二和第三个时间点之间。返工则是变更请求之后的一切工作,包括作者的修复、重新审查以及任何后续的往返沟通。

Instrument those four per PR, split by provenance and by repository, and the tool decision falls out of the data instead of the demo. If wait dominates, the lever is the trigger. If rework dominates, the lever is whether the tool produces one precise change request or a stream of comments the author answers in four passes. 针对每个 PR 测量这四个指标,并按来源和代码库进行拆分,那么工具的选择将基于数据而非演示。如果等待时间占主导,关键在于触发机制。如果返工占主导,关键在于工具是产生一个精确的变更请求,还是产生一连串需要作者分四次回复的评论。

If reading time dominates and your diffs are large, reading speed is finally the right thing to buy. The rework share is where vendor differences show up most clearly. Copilot PRs drew the most reviews and change requests in the Kraishan dataset, which means more round trips per change, not slower readers. 如果阅读时间占主导且你的代码差异很大,那么提升阅读速度才是正确的投资方向。返工占比是供应商差异体现最明显的地方。在 Kraishan 的数据集中,Copilot 的 PR 引发的审查和变更请求最多,这意味着每个变更需要更多的往返沟通,而不是因为阅读者变慢了。

Round trips are also the cost that a review tool can multiply. Ten findings in one pass is one loop. The same ten findings spread across four review cycles is four loops, and each loop carries a context reload, a re-review, and a chance that the author argues with a comment that already stopped being true. 往返沟通也是审查工具可能导致成本倍增的地方。一次性发现十个问题是一个循环。如果同样的十个问题分散在四次审查周期中,那就是四个循环,每个循环都伴随着上下文重新加载、重新审查,以及作者与一条已经过时的评论进行争论的风险。

Trigger timing is a review feature. When a review runs matters as much as what it says. A review that fires on PR open puts its feedback into the same window as the wait, so the author and the assigned reviewer see the same first pass. A review that fires after CI, or only when a human asks for it, adds its latency on top of the queue instead of overlapping it. 触发时机是审查功能的一部分。审查何时运行与审查内容同样重要。在 PR 开启时触发的审查将反馈放入等待窗口中,这样作者和指定的审查者能看到相同的初次反馈。而在 CI 完成后或仅在人工请求时才触发的审查,则是在排队时间之上增加了延迟,而非与之重叠。

Ask any vendor for the trigger list: open, ready-for-review, CI green, manual comment command. Where the tool sits in that list determines whether it can touch the biggest number in the cycle. 询问任何供应商其触发列表:开启、准备审查、CI 通过、手动评论命令。工具在列表中的位置决定了它是否能触及周期中占比最大的那个数字。

Rules cut round trips only when they fire on the right files. Stack Overflow’s post on coding guidelines for AI and people contains the line that ties standards to cycle time. Code review will be most engineers’ first look at code they did not write. Heroku chief architect Vish Abrams makes the related point there that principles seasoned engineers assume, like DRY, are not common knowledge to an agent. 规则只有在正确的文件上触发时才能减少往返沟通。Stack Overflow 关于 AI 与人类编码准则的文章中提到,标准与周期时间息息相关。代码审查将是大多数工程师第一次看到他们未编写的代码。Heroku 首席架构师 Vish Abrams 指出,资深工程师认为理所当然的原则(如 DRY),对智能体来说并非共识。

Rules are supposed to move the standards check earlier so a human reviewer does not spend the first pass on naming and layout, which is exactly the work that generates change requests when it is missed. A rule only reduces that work if it fires on the right files. A payments rule that fires on the CLI tool produces comments the author has to triage, and triage is a round trip. 规则旨在将标准检查提前,这样人类审查者就不必在第一轮审查中纠结于命名和布局——而这正是如果被忽略就会产生变更请求的工作。规则只有在正确的文件上触发才能减少这种工作。如果一个支付规则在 CLI 工具上触发,就会产生作者必须处理的评论,而处理这些评论就是一次往返沟通。

Kodus imports the rule files teams already keep, including AGENTS.md, CLAUDE.md, .cursorrules, Copilot instruction files, Windsurf rules, and docs/coding-standards, scopes each rule to a path glob, and discovers nested files so a services/billing/CLAUDE.md applies to services/billing/** without extra setup. Kodus 会导入团队现有的规则文件,包括 AGENTS.md、CLAUDE.md、.cursorrules、Copilot 指令文件、Windsurf 规则以及 docs/coding-standards,将每条规则限定在路径通配符内,并自动发现嵌套文件,因此 services/billing/CLAUDE.md 无需额外设置即可应用于 services/billing/**。

On self-hosted deployments it writes a per-file trace to the API log under the marker [kody-rules-eval], listing the rule ids selected into the prompt for each reviewed file. The same page documents a limit worth knowing before you plan around it: an unlicensed Community Edition instance evaluates at most 10 rules per review, oldest first, so a rule past the tenth may not fire at all. 在自托管部署中,它会将每个文件的追踪信息写入 API 日志,标记为 [kody-rules-eval],列出每个被审查文件在提示词中选定的规则 ID。同一页面还记录了一个在规划前值得了解的限制:未授权的社区版实例每次审查最多评估 10 条规则(按时间顺序),因此第 10 条之后的规则可能根本不会触发。

A rule that silently does not run is one fewer thing checked before a human, and one more thing that turns up later as a comment. The trace matters beyond debugging. It is the difference between a tool that states your standards load correctly and one where you can point at the rule id and the file it evaluated. Everything else in this section assumes that mechanism works. 一条静默失效的规则意味着在人类介入前少了一项检查,也意味着未来多了一条评论。追踪功能的作用不仅限于调试。它区分了一个仅仅声称“标准已正确加载”的工具,和一个你可以明确指出规则 ID 及评估文件的工具。本节其余部分均假设该机制正常运作。

What each tool documents about the rework half. The criteria below are the ones that affect r… 各工具关于返工部分的文档说明。以下标准是影响返工的因素……