The complement of true is true, except when it's false

The complement of true is true, except when it’s false

真值的补码是真值,除非它是假值

I recently looked at P4313R1, a standards proposal paper which adds a set of bitmask operations for enums, using a C++26 annotation to opt-in. The core idea being that, to use an example from the paper, given code like the below: 最近我研究了 P4313R1,这是一份标准提案,旨在通过 C++26 注解为枚举类型添加一组位掩码(bitmask)操作。其核心思想是,引用提案中的一个例子,给定如下代码:

enum class [[=std::bitmask_type]] Permission {
    None = 0,
    Read = 1 << 0,
    Write = 1 << 1,
    Execute = 1 << 2,
};

The [[=std::bitmask_type]] annotation would automatically imbue Permission with an accessible set of bitwise operations so that you as a user needn’t write them out yourself. This got me thinking about the murky underbelly of C++ integer operations. Your C++ compiler will happily compute bitwise operations for any integral type as a builtin operation. This extends to types which we don’t traditionally think of as integers, such as wchar_t, the UTF character types char8_t through char32_t, and bool. [[=std::bitmask_type]] 注解会自动为 Permission 赋予一组可用的位运算,这样作为用户,你就不必自己编写这些代码了。这让我开始思考 C++ 整数运算中那些模糊的底层细节。C++ 编译器乐于将任何整型作为内置操作进行位运算。这甚至扩展到了我们传统上不认为是整数的类型,例如 wchar_t、从 char8_t 到 char32_t 的 UTF 字符类型,以及 bool。

But slightly more happens here than meets the eye. Because when you attempt to perform an operation like a | b, and the type of a and b is an integral type smaller than int, the language does not operate on the bit patterns of a and b directly. They undergo integral promotion - they are promoted up to int as an intermediate state, the bit patterns of these two ints are combined, and the result is returned to you, still as an int. 但这里发生的事情比表面看起来要复杂一些。因为当你尝试执行类似 a | b 的操作,且 a 和 b 的类型是小于 int 的整型时,语言并不会直接对 a 和 b 的位模式进行操作。它们会经历“整型提升”(integral promotion)——它们被提升为 int 作为中间状态,这两个 int 的位模式进行合并,然后将结果返回给你,结果仍然是一个 int。

Consider the below: 请看下例:

// Two shorts
constexpr short perm_A {1 << 0};
constexpr short perm_B {1 << 1};

// And the result type of running a bitwise operation on them is int
static_assert(std::same_as<decltype(perm_A | perm_B), int>);

Now, for the most part this is harmless - if you cast the above perm_A | perm_B back to short then the unnecessary bytes are truncated away and you’re left with a short which contains exactly the value it would have held if you’d combined perm_A and perm_B as short directly. And if you want to be certain that you are keeping your types consistent, and avoid the pernicious bugs of silent narrowing conversions, you can get into the habit of static_cast-ing the result of your bitwise operation to its original type. The important thing here however is that integral promotion is not optional. Unlike most other areas of C++ the developer doesn’t get a choice - your types will unavoidably be promoted for these operations. 在大多数情况下,这并无大碍——如果你将上述 perm_A | perm_B 强制转换回 short,多余的字节会被截断,剩下的 short 正好包含你直接以 short 组合 perm_A 和 perm_B 时应有的值。如果你想确保类型一致,并避免隐式窄化转换带来的隐蔽 Bug,你可以养成将位运算结果 static_cast 回原始类型的习惯。然而,这里重要的一点是,整型提升不是可选的。与 C++ 的大多数其他领域不同,开发者没有选择权——在这些操作中,你的类型不可避免地会被提升。

The second part of what makes this a hazard for bool specifically is boolean conversion, a special case which does not truncate but instead explicitly converts zero to false and non-zero values to true. For these non-zero values, whatever bit pattern was previously stored is discarded and replaced with true, which has an integer value of 1. 对于 bool 类型而言,使其成为隐患的第二个因素是布尔转换(boolean conversion)。这是一个特殊情况,它不会进行截断,而是明确地将零转换为 false,将非零值转换为 true。对于这些非零值,之前存储的任何位模式都会被丢弃,并替换为 true(其整数值为 1)。

So if we take code like this: int x{10}; bool b{static_cast<bool>(x)}; and look at the generated asm (in this case from unoptimised x86-64 gcc 16.2): 因此,如果我们采用如下代码:int x{10}; bool b{static_cast<bool>(x)}; 并查看生成的汇编代码(此处为未优化的 x86-64 gcc 16.2):

mov DWORD PTR [rbp-4], 10    ; Store the value of 10
cmp DWORD PTR [rbp-4], 0     ; Then compare to 0, set ZF if x is zero
setne al                     ; Write 1 to AL if ZF is clear
mov BYTE PTR [rbp-5], al     ; Store the result

The standard (specifically [conv.bool]) bases converting to bool on the only possible values which bool can hold - true and false, regardless of whatever bit pattern may have originally been used to create them. 标准(具体为 [conv.bool])将转换为 bool 的依据建立在 bool 仅能持有的两个值——true 和 false 之上,而无论最初用于创建它们的位模式是什么。

Putting it together

综合分析

With all of that covered, let’s talk through what happens when you try to evaluate ~true and then cast the result to a bool: 了解了以上内容,让我们来看看当你尝试计算 ~true 并将结果转换为 bool 时会发生什么:

  1. The value true is promoted to an int with a value of 1.
  2. 值 true 被提升为值为 1 的 int。
  3. The complement of 1 as an int is calculated as -2, as two’s complement behaviour is required as of C++20.
  4. 1 作为 int 的补码计算结果为 -2(自 C++20 起,强制要求使用补码行为)。
  5. The value of -2 is then cast back down to bool, undergoes boolean conversion, and since -2 is non-zero, becomes true.
  6. 值 -2 随后被强制转换回 bool,经历布尔转换,由于 -2 是非零值,因此变为 true。

There you have it, the complement of true is true, or spelled in C++ static_cast<bool>(~true) == true. 结果就是,true 的补码是 true,用 C++ 代码表示即 static_cast<bool>(~true) == true。

This brings us back to enums. An enum is permitted to use any integral type as its underlying type, including our good friend bool. So let’s define one: 这让我们回到了枚举类型。枚举允许使用任何整型作为其底层类型,包括我们的老朋友 bool。那么让我们定义一个:

enum class [[=std::bitmask_type]] boolean : bool {
    FALSE,
    TRUE,
};

Let’s also look at the bitwise complement operator as laid out in P4313R1: 再来看看 P4313R1 中定义的按位取反运算符:

template<bitmask-like T>
constexpr T operator~ (T lhs) noexcept {
    return static_cast<T>(~to_underlying(lhs));
}

By now the workings of that operator should be a familiar shape - first we convert the enum to its underlying type (in our case bool), then integral promotion is applied if that type is lower ranked than int, then we perform the bitwise operation, then we cast back to the enum type. As we would expect: static_assert(~boolean::TRUE == boolean::TRUE); 到目前为止,该运算符的工作方式应该很熟悉了——首先我们将枚举转换为其底层类型(在本例中为 bool),如果该类型的秩低于 int,则应用整型提升,然后执行位运算,最后强制转换回枚举类型。正如我们所预期的:static_assert(~boolean::TRUE == boolean::TRUE);

This has all the potential to be a slightly confusing corner case in the language. 这完全有可能成为语言中一个令人困惑的边缘情况。

When it’s false

当它是假值时

There is one exceptional case here which compounds the problem. If we try to run that same example in gcc, we get a different result: static_assert(~boolean::TRUE == boolean::FALSE); 这里有一个特殊情况加剧了这个问题。如果我们尝试在 gcc 中运行同样的例子,会得到不同的结果:static_assert(~boolean::TRUE == boolean::FALSE);

So what is going on with gcc? If we move the operation to runtime by defining two functions which depend on it as a runtime value… 那么 gcc 到底发生了什么?如果我们通过定义两个依赖于运行时值的函数将操作移至运行时……

What we see is that when the type is not spelled bool, gcc will truncate down to the lowest bit rather than perform a boolean conversion. Any even value, when cast down, will produce boolean::FALSE, and any odd one will produce boolean::TRUE. This is unique to enum. 我们看到,当类型不是显式的 bool 时,gcc 会截断到最低位,而不是执行布尔转换。任何偶数值在强制转换后都会产生 boolean::FALSE,而任何奇数值都会产生 boolean::TRUE。这是枚举类型特有的现象。