Rust's derive often implies inline

Rust’s derive often implies inline

在 Rust 中,实现核心 trait(如 Debug、Display 和 Clone)最常见的方法之一是使用 #[derive(...)],例如:

#[derive(Debug)]
struct Widgets {
    foo: u32,
    bar: usize,
}

What I didn’t know until recently is that Rust currently emits #[inline] as part of these derivations. This is seemingly not guaranteed, but is implied by example in the reference and can also be seen if one expands the macros. Using the example above, this is what you get when you expand the #[derive(Debug)] in the playground:

直到最近我才知道,Rust 目前在这些派生实现中会自动加上 #[inline]。这似乎并没有被明确保证,但在参考文档的示例中有所暗示,并且通过展开宏也可以观察到。使用上面的例子,在 playground 中展开 #[derive(Debug)] 后,你会得到:

struct Widgets {
    foo: u32,
    bar: usize,
}

#[automatically_derived]
impl ::core::fmt::Debug for Widgets {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field2_finish(f, "Widgets", "foo", &self.foo, "bar", &&self.bar)
    }
}

This is almost always what we want: #[inline] is just a hint, and typical derived Debug, Clone, etc. implementations benefit from being inlined (since they’re often trivial). But not always! Imagine an error hierarchy like this:

这几乎总是我们想要的结果:#[inline] 只是一个提示,典型的派生 Debug、Clone 等实现通过内联会受益(因为它们通常很琐碎)。但并非总是如此!想象一下这样的错误层级:

#[derive(Debug)]
struct ErrorA {
    lots: String,
    of: String,
    chunky: String,
    fields: String,
    within: String,
    this: String,
    r#type: String,
}

#[derive(Debug)]
struct ErrorB { inner: ErrorA }

#[derive(Debug)]
struct ErrorC { inner: ErrorB }

#[derive(Debug)]
enum Errors {
    A(ErrorA),
    B(ErrorB),
    C(ErrorC),
}

(展开后的代码如下:)

struct ErrorA { lots: String, of: String, chunky: String, fields: String, within: String, this: String, r#type: String, }

#[automatically_derived]
impl ::core::fmt::Debug for ErrorA {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        let names: &'static _ = &["lots", "of", "chunky", "fields", "within", "this", "type"];
        let values: &[&dyn ::core::fmt::Debug] = &[&self.lots, &self.of, &self.chunky, &self.fields, &self.within, &self.this, &&self.r#type];
        ::core::fmt::Formatter::debug_struct_fields_finish(f, "ErrorA", names, values)
    }
}

#[automatically_derived]
impl ::core::default::Default for ErrorA {
    #[inline]
    fn default() -> Self {
        Self { lots: ::core::default::Default::default(), of: ::core::default::Default::default(), chunky: ::core::default::Default::default(), fields: ::core::default::Default::default(), within: ::core::default::Default::default(), this: ::core::default::Default::default(), r#type: ::core::default::Default::default(), }
    }
}

struct ErrorB { inner: ErrorA, }

#[automatically_derived]
impl ::core::fmt::Debug for ErrorB {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field1_finish(f, "ErrorB", "inner", &&self.inner)
    }
}

struct ErrorC { inner: ErrorB, }

#[automatically_derived]
impl ::core::fmt::Debug for ErrorC {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        ::core::fmt::Formatter::debug_struct_field1_finish(f, "ErrorC", "inner", &&self.inner)
    }
}

enum Errors { A(ErrorA), B(ErrorB), C(ErrorC), }

#[automatically_derived]
impl ::core::fmt::Debug for Errors {
    #[inline]
    fn fmt(&self, f: &mut ::core::fmt::Formatter) -> ::core::fmt::Result {
        match self {
            Self::A(__self_0) => ::core::fmt::Formatter::debug_tuple_field1_finish(f, "A", &__self_0),
            Self::B(__self_0) => ::core::fmt::Formatter::debug_tuple_field1_finish(f, "B", &__self_0),
            Self::C(__self_0) => ::core::fmt::Formatter::debug_tuple_field1_finish(f, "C", &__self_0),
        }
    }
}

That’s a lot of code that can get inlined for each invocation of the Debug implementation of Errors, which can occur repeatedly in e.g. debug or trace logging. In fact, it’s so much code that it can turn out to be a non-trivial amount of a Rust binary’s total size: we found that we could shrink uv’s binary size by approximately 160KB by preventing rustc from inlining a given Debug implementation.

对于 Errors 的 Debug 实现,每次调用时都会内联大量的代码,这种情况在调试或追踪日志中可能会反复发生。事实上,这些代码量非常大,以至于会占据 Rust 二进制文件总大小中不可忽视的一部分:我们发现,通过阻止 rustc 内联特定的 Debug 实现,我们可以将 uv 的二进制文件大小缩小约 160KB。

This was surprising to me on two levels: the size cost added up fast, and rustc (seemingly) did not apply a limit to the size or number of times a Debug implementation was inlined. I suspect this is the right decision in many programs, however! This hierarchy is drastically simplified: real world Rust applications often have deeply nested error enumerations with nontrivial numbers of fields.

这让我感到惊讶,原因有二:一是体积成本增加得很快,二是 rustc(似乎)没有对 Debug 实现的内联大小或次数设置限制。不过,我怀疑这对于许多程序来说是正确的决定!上述层级结构被极大地简化了:现实世界中的 Rust 应用程序通常具有深度嵌套的错误枚举,且包含大量字段。