How to speed up the Rust compiler in September 2026
How to speed up the Rust compiler in September 2026
如何在 2026 年 9 月加速 Rust 编译器
My last post on the Rust compiler’s performance was two months ago and a lot has happened since then. 我上一篇关于 Rust 编译器性能的文章是在两个月前发布的,从那时起发生了很多变化。
Overall progress
整体进展
The measurements for the period 2026-07-29 to 2026-09-28 can be seen here. The mean wall-time reduction was 4.57%, which is a remarkable improvement in just two months. Of the 629 benchmark measurements, 555 of them improved and only 74 regressed. A number of benchmarks saw double-digit percentage reductions. The technical term for this result is “a sea of green”. 2026 年 7 月 29 日至 2026 年 9 月 28 日期间的测量结果可以在这里查看。平均挂钟时间(wall-time)减少了 4.57%,这在短短两个月内是一个显著的进步。在 629 项基准测试中,有 555 项得到了改善,只有 74 项出现了倒退。许多基准测试的性能提升达到了两位数百分比。这种结果的专业术语被称为“一片绿海”(a sea of green)。
rustdoc
rustdoc
In my last post I mentioned how Noah Lev got some enormous speed wins on rustdoc. He recently wrote a post explaining in some detail exactly how he did this. It’s an interesting and satisfying read. 在上一篇文章中,我提到了 Noah Lev 如何在 rustdoc 上取得了巨大的速度提升。他最近写了一篇文章,详细解释了他具体是如何做到的。这是一篇既有趣又令人满意的文章。
Clippy
Clippy
#159642: In this PR Jakub Beránek enabled PGO for Clippy, giving wall-time improvements across most Clippy benchmarks, in the best case by 18%! #159642:在此 PR 中,Jakub Beránek 为 Clippy 启用了 PGO(配置引导优化),使得大多数 Clippy 基准测试的挂钟时间得到了改善,最好情况下提升了 18%!
LLVM update
LLVM 更新
#158734: In this PR Nikita Popov upgraded the LLVM version used by the compiler to LLVM 23. As often happens when we upgrade LLVM, we saw some nice speedups. The mean wall-time reduction across all benchmarks was 1.2%, which might not sound like much but is really impressive for a single PR. Great work from the LLVM folks! #158734:在此 PR 中,Nikita Popov 将编译器使用的 LLVM 版本升级到了 LLVM 23。正如我们升级 LLVM 时经常发生的那样,我们看到了一些不错的速度提升。所有基准测试的平均挂钟时间减少了 1.2%,这听起来可能不多,但对于单个 PR 来说确实令人印象深刻。LLVM 团队干得漂亮!
The new borrow checker
新的借用检查器
The new borrow checker, Polonius Alpha (no relation to Napoleon Dynamite), was enabled on Nightly. It is more precise than the existing borrow checker and accepts some valid programs that the old borrow checker would reject. It does do more work than the old borrow checker, enough to make a measurable difference to compile time in a minority of cases, including the popular serde crate. Fortunately, Jack Huey has been on the case. 新的借用检查器 Polonius Alpha(与《拿破仑炸药》无关)已在 Nightly 版本中启用。它比现有的借用检查器更精确,并能接受一些旧检查器会拒绝的有效程序。它的工作量确实比旧检查器大,足以在少数情况下(包括流行的 serde crate)对编译时间产生可测量的影响。幸运的是,Jack Huey 一直在跟进这个问题。
#161938: In this PR Jack made some liveness computations lazy, which reduced instruction counts for serde by 3-5%, and for some other benchmarks by less than 1%. #161938:在此 PR 中,Jack 将一些活跃度计算改为惰性计算,这使得 serde 的指令数减少了 3-5%,其他一些基准测试减少了不到 1%。
#163027: In this PR Jack adjusted a data structure and tweaked some inlining, for mostly sub-1% instruction count reductions across numerous benchmarks. There is more work to be done to reduce the remaining Polonius Alpha regressions, but it’s worth noting that the “sea of green” shows these regressions were swamped by the many other recent improvements. #163027:在此 PR 中,Jack 调整了一个数据结构并微调了一些内联,使得众多基准测试的指令数减少了不到 1%。要减少剩余的 Polonius Alpha 性能倒退还有更多工作要做,但值得注意的是,“一片绿海”表明这些倒退已被最近的许多其他改进所淹没。
The new trait solver
新的特征求解器
The new trait solver, Penelope Hammertime, [Ed. note: is that right?] was also enabled on Nightly. As I said, a lot has been happening. Like the new borrow checker, the new trait solver is slower in a minority of cases. Jana Dönszelmann wrote a detailed post about the efforts to improve the performance of this new solver. Jana’s post is detailed enough that I won’t say much more about the large amount of ongoing work on the new solver, but I will mention in passing the PRs I made: #160479, #160605, #160801, #160892, #161077, and #161211. Some of these reduced compile times greatly for certain outlier crates: 50% here, 25% there, 15% there, and even more on one stress test. And I am not the only one who has made progress here… go read Jana’s post. 新的特征求解器 Penelope Hammertime [编者注:名字对吗?] 也已在 Nightly 版本中启用。正如我所说,最近发生了很多事情。和新的借用检查器一样,新的特征求解器在少数情况下会变慢。Jana Dönszelmann 写了一篇详细的文章,介绍了改进该求解器性能的努力。Jana 的文章非常详尽,所以我不会再多说关于新求解器正在进行的大量工作,但我会顺便提一下我所做的 PR:#160479, #160605, #160801, #160892, #161077 和 #161211。其中一些极大地缩短了某些异常 crate 的编译时间:有的缩短了 50%,有的 25%,有的 15%,在一次压力测试中甚至更多。而且我不是唯一在这里取得进展的人……去读读 Jana 的文章吧。
xmakro
xmakro
New contributor xmakro continued their run of good improvements. 新贡献者 xmakro 继续保持着良好的改进势头。
#157281: In this PR xmakro optimized impl handling when building the specialization graph. This gave a mean cycle count reduction of 1.58% across all benchmarks, which is huge for a single PR. #157281:在此 PR 中,xmakro 在构建特化图(specialization graph)时优化了 impl 处理。这使得所有基准测试的平均周期数减少了 1.58%,对于单个 PR 来说这非常巨大。
#158059: In this PR xmakro optimized one aspect of the loading of incremental compilation data, reducing instruction counts across multiple benchmarks, in the best case by 6%. #158059:在此 PR 中,xmakro 优化了增量编译数据加载的一个方面,减少了多个基准测试的指令数,最好情况下减少了 6%。
#160473: In this PR xmakro avoided some allocations in a hot obligations processing path, reducing instruction counts across numerous benchmarks, in the best case by 2%. #160473:在此 PR 中,xmakro 避免了热点义务处理路径中的一些内存分配,减少了众多基准测试的指令数,最好情况下减少了 2%。
#160268: In this PR xmakro avoided a lot of allocations by changing the old/new trait solver selection code to use static dispatch instead of dynamic dispatch. This gave mostly sub-1% instruction count reductions across a number of benchmarks. This hot allocation path had been showing up in profiles for a while and I had earlier tried exactly the same idea in #155714. But I got regressions on a couple of benchmarks, possibly due to slightly different choices of where to place some #[inline] attributes. It was good to see this obvious inefficiency fixed. #160268:在此 PR 中,xmakro 通过将旧/新特征求解器的选择代码改为使用静态分发而非动态分发,避免了大量内存分配。这使得多个基准测试的指令数减少了不到 1%。这个热点分配路径在性能分析中已经出现了一段时间,我之前在 #155714 中也尝试过完全相同的想法。但我当时在几个基准测试中遇到了性能倒退,这可能是由于在放置 #[inline] 属性的位置上选择了略有不同。很高兴看到这个明显的低效问题得到了解决。
Dataflow analysis
数据流分析
#160193: In this PR I changed the CFG traversal algorithm used by the dataflow analyses in the compiler. These analyses iterate to a fixpoint and the traversal algorithm can affect how quickly the fixpoint is reached. For most code the new algorithm makes no difference, but the cranelift-codegen crate has one enormous function with over 18,000 basic blocks. The old algorithm required 1.5 million calls to apply_effects_in_block to reach a fixpoint for the EverInitializedPlaces analysis used by the borrow checker; the new algorithm requires 90,000. This gave an enormous ~30% wall-time reduction for a check build of this crate. #160193:在此 PR 中,我更改了编译器中数据流分析所使用的 CFG 遍历算法。这些分析会迭代直到达到不动点(fixpoint),而遍历算法会影响达到不动点的速度。对于大多数代码,新算法没有区别,但 cranelift-codegen crate 有一个包含超过 18,000 个基本块的巨大函数。旧算法需要 150 万次 apply_effects_in_block 调用才能为借用检查器使用的 EverInitializedPlaces 分析达到不动点;而新算法只需要 90,000 次。这使得该 crate 的检查构建(check build)挂钟时间减少了约 30%。
#160033: In this PR I made EverInitializedPlaces more efficient again, this time by not tracking unnecessary data for projections. This reduced instruction counts on the match-stress benchmark by 17%, and on a few other benchmarks by less than 1%. #160033:在此 PR 中,我再次提高了 EverInitializedPlaces 的效率,这次是通过不跟踪投影(projections)的不必要数据。这使得 match-stress 基准测试的指令数减少了 17%,其他一些基准测试减少了不到 1%。
LLMs
大语言模型
They’ve gotten very good at certain kinds of analysis. I’m still writing all my own code and text, because (a) that’s paramount, and (b) the project policy requires it, but I had useful LLM analysis assistance on several of the PRs mentioned in this post. Anyway, enough about that. 它们在某些类型的分析上已经变得非常出色。我仍然坚持自己编写所有的代码和文本,因为 (a) 这至关重要,(b) 项目政策有此要求,但我确实在本文提到的几个 PR 中获得了有用的 LLM 分析辅助。总之,关于这一点就说这么多。
Miscellaneous
其他
#160535: In this PR Chris Denton increased the default stack size used by the compiler, which allowed the removal of ensure_sufficient_stack, a manual stack extension mechanism sprinkled about in places prone to high levels of recursion. There was a lot of discussion about this one because it can be difficult to decide how to best deal with stack exhaustion. But the performance effects are clear, with reduced instruction counts across many benchmarks, in the best case by almost 3%. #160535:在此 PR 中,Chris Denton 增加了编译器使用的默认栈大小,这使得可以移除 ensure_sufficient_stack——这是一种手动栈扩展机制,散布在容易出现高递归的地方。关于这一点有很多讨论,因为很难决定如何最好地处理栈溢出。但性能效果很明显,许多基准测试的指令数都有所减少,最好情况下减少了近 3%。
#160506: The project uses a lot of “rollup” PRs, where multiple PRs are merged together. This is because we don’t have sufficient CI capacity to merge every PR individually. Normally PRs that affect perform… #160506:该项目使用了许多“汇总”(rollup)PR,即将多个 PR 合并在一起。这是因为我们没有足够的 CI 能力来单独合并每个 PR。通常,影响性能的 PR…