Typeclasses vs Modules

Typeclasses vs Modules

For a long time I’ve noticed that there’s a lot of persistent confusion about the differences and similarities between module systems (as seen in languages like OCaml) and typeclasses (as seen in languages like Haskell or Rust). This is an attempt to clear up the confusion. 长期以来,我注意到人们对于模块系统(如 OCaml 语言中所见)与类型类(如 Haskell 或 Rust 语言中所见)之间的异同存在着持续的困惑。本文旨在厘清这些混淆。

The short answer

简短的回答

The main goal of typeclasses as a language construct is to provide ad-hoc polymorphism. This is also known as operator overloading or type-based dispatch. The main point of ad-hoc polymorphism is as a convenience feature for programming in the small: it lets you reuse the same identifiers across types. For example, I might want to use the + symbol to denote both addition on integers and addition on vectors of floats (Importantly, the vector type might be defined in a random library, so we can’t hardcode this behaviour in the language). 作为一种语言结构,类型类的主要目标是提供特设多态(ad-hoc polymorphism)。这也被称为运算符重载或基于类型的分发。特设多态的主要意义在于作为一种“小规模编程”的便利特性:它允许你在不同类型间复用相同的标识符。例如,我可能希望使用 + 符号来同时表示整数加法和浮点向量加法(重要的是,向量类型可能定义在某个随机库中,因此我们无法在语言层面硬编码这种行为)。

The main goal of module systems as a language construct is to provide modular abstraction. Some parts of this feature have names in the broader software industry, like dependency injection, encapsulation, and information hiding. The main point of modular abstraction is to enable more efficient programming in the large, by allowing you to express your program decomposition explicitly. It gives you a compositional way to break down programs into component pieces recursively. 作为一种语言结构,模块系统的主要目标是提供模块化抽象。这一特性的某些部分在更广泛的软件行业中有对应的名称,如依赖注入、封装和信息隐藏。模块化抽象的主要意义在于通过允许显式表达程序分解,从而实现更高效的“大规模编程”。它提供了一种组合式的方法,将程序递归地拆解为各个组件。

The main factors that I think lead to confusion are: historically, programming languages have tended to pick one or the other; they share some of the underlying mechanism (a way to specify an interface and a way to show that some bundle of code conforms to the interface); you sometimes can kinda sorta emulate one using the other (i.e use typeclasses to provide modular abstraction, and use modules to provide ad-hoc polymorphism). However, while emulation is possible in many cases it’s usually not optimal. 我认为导致混淆的主要因素有:从历史上看,编程语言往往倾向于二选一;它们共享一些底层机制(指定接口的方式,以及证明某段代码符合该接口的方式);有时你可以在某种程度上用一种机制模拟另一种(例如,使用类型类提供模块化抽象,或使用模块提供特设多态)。然而,虽然在许多情况下模拟是可行的,但通常并不是最优解。

Why and What: Typeclasses

缘由与定义:类型类

The motivational idea behind typeclasses is basically, if you’re going to overload a symbol, it should “morally” mean the same thing. For example the symbol + denotes the abstract idea of addition, and we can have different implementations for different types. This is in contrast to for example C++ where << can mean bitshift or stream redirection depending on the type, which have nothing in common. The goal here is for the meaning of programs to be easier to understand (relative to the unprincipled C++ thing), while reducing cognitive overhead by reducing the number of symbols. We can learn a few symbols that work with abstract structures, like + or >>=, and use them across a variety of different concrete structures. 类型类背后的动机理念基本上是:如果你要重载一个符号,它在“逻辑上”应该具有相同的含义。例如,符号 + 表示加法的抽象概念,我们可以针对不同类型有不同的实现。这与 C++ 形成了对比,在 C++ 中 << 可能表示位移或流重定向,具体取决于类型,而这两者毫无共同之处。这里的目标是使程序的含义更易于理解(相对于 C++ 那种无原则的做法),同时通过减少符号数量来降低认知负担。我们可以学习少量适用于抽象结构的符号(如 + 或 >>=),并将它们应用于各种不同的具体结构中。

If we want to be a bit more serious about this we need to consider the semantics as well. How do we make sure + does in fact denote an abstract addition operation? Well, we can add laws to the typeclass. For instance we can say that addition has to satisfy (a+b)+c=a+(b+c) (though this isn’t true for floats), and then check that new implementations do in fact satisfy the law using a property test or a proof witness. In general, each typeclass can be bundled with some description of semantics, ideally machine checked in the form of a test suite or proof obligation. 如果我们想更严谨地对待这一点,还需要考虑语义。我们如何确保 + 确实表示抽象的加法运算?我们可以为类型类添加定律。例如,我们可以规定加法必须满足 (a+b)+c = a+(b+c)(尽管这对浮点数不成立),然后使用属性测试或证明见证来检查新的实现是否确实满足该定律。通常,每个类型类都可以捆绑一些语义描述,理想情况下是以测试套件或证明义务的形式进行机器验证。

(Note: The provided Haskell code snippet and the “Why and What: Modules” section follow the same pattern of alternating translation.) (注:原文后续的 Haskell 代码片段及“缘由与定义:模块”部分亦应遵循此交替翻译格式。)