Iroh global content discovery

Iroh 全球内容发现

What got me excited about IPFS many years ago, briefly after it was announced, was being able to publish a personal website, blog post, or political pamphlet and have it remain available globally as long as enough people are interested in the content. As governments have been more sophisticated in their firewalling methods, this use case for circumvention tools and permissionless global content discovery is still incredibly relevant. Recent events have added some urgency to this. IPFS shipyard is shutting down. This does not mean that IPFS will stop working, but it does not bode well for the future of the project.

多年前,在 IPFS 刚发布不久时,它让我感到兴奋的是:我能够发布个人网站、博客文章或政治宣传册,只要有足够多的人对这些内容感兴趣,它们就能在全球范围内保持可用。随着各国政府在防火墙技术上变得越来越复杂,这种用于规避审查工具和无需许可的全球内容发现的使用场景,依然具有极高的现实意义。最近发生的一些事件增加了这种紧迫感——IPFS Shipyard 正在关闭。这并不意味着 IPFS 将停止工作,但这对于该项目的未来来说并不是一个好兆头。

The state of the art

技术现状

When we had to solve hole punching, we started looking at existing systems and chose the best open source system as an initial starting point for our own implementation. So let’s do the same for global content discovery. There are a number of projects trying to solve this problem. But one project stands above all others: BitTorrent. It just works and has done so for over two decades. So let’s take a look at what makes BitTorrent the current leader in permissionless global content discovery.

当我们必须解决“打洞”(hole punching)问题时,我们开始研究现有的系统,并选择了最优秀的开源系统作为我们自己实现方案的起点。那么,让我们以同样的方式来处理全球内容发现问题。目前有许多项目试图解决这个问题,但有一个项目脱颖而出:BitTorrent。它不仅运行良好,而且已经稳定运行了二十多年。因此,让我们来看看是什么让 BitTorrent 成为目前无需许可的全球内容发现领域的领导者。

BitTorrent has a relatively simple protocol for blob transfer and a DHT called Mainline for global content discovery. The transfer protocol and content discovery are separate systems. In fact, the DHT was developed later than the transfer protocol. BitTorrent was released in 2001 using centralized trackers for content discovery; the Mainline DHT was added in 2005.

BitTorrent 拥有一个相对简单的二进制大对象(blob)传输协议,以及一个名为 Mainline 的 DHT(分布式哈希表)用于全球内容发现。传输协议和内容发现是两个独立的系统。事实上,DHT 的开发时间晚于传输协议。BitTorrent 于 2001 年发布时使用中心化追踪器(trackers)进行内容发现;Mainline DHT 则是在 2005 年才加入的。

Transfer protocol

传输协议

BitTorrent works by creating a .torrent file that contains information about the data to be downloaded. The file is encoded using bencode, which is conceptually similar to JSON.

BitTorrent 的工作原理是创建一个包含待下载数据信息的 .torrent 文件。该文件使用 bencode 编码,其概念类似于 JSON。

{
  "announce": "http://tracker.example.com/announce",
  "comment": "Ubuntu 24.04 desktop image",
  "created by": "mktorrent 1.1",
  "creation date": 1714521600,
  "info": {
    "name": "ubuntu-24.04-desktop-amd64.iso",
    "piece length": 262144,
    "length": 5820411904,
    "pieces": "9358b5d71f8007ac926000961ad79d865d9652bc 4cce840457bcfac1dd67d864e9386422eee70904 … many more 20-byte hashes …"
  }
}

pieces contains the concatenated SHA-1 hashes of all piece length-sized pieces. The transfer protocol downloads blocks of these pieces from different peers. The transfer protocol allows requests for ranges within pieces, but you can only validate a piece after you have downloaded it completely and computed its SHA-1 hash.

pieces 字段包含了所有固定长度数据块的 SHA-1 哈希值的拼接。传输协议会从不同的节点(peers)下载这些数据块。虽然传输协议允许请求数据块内的特定范围,但你必须在完整下载该数据块并计算出其 SHA-1 哈希值后,才能对其进行验证。

I don’t want to dwell on this too long, but I think that BLAKE3 verified streaming is a superior replacement for the transfer protocol. We implemented a protocol called iroh-blobs that uses BLAKE3 verified streaming for sharing blobs of data over iroh connections. With BLAKE3, we only need a single root hash, and neither a piece length parameter nor hashes of the individual pieces. This is what allows iroh-blobs to validate any piece of a large blob without needing intermediate hashes. For a visual explanation of how it works, see my BLAKE3 and Bao deep dive, which covers verified streaming, outboard encoding, and range requests. We still have some work to do to make multiprovider downloads of single blobs efficient, but the streaming protocol itself with its fine grained validation is superior to the BitTorrent transfer protocol.

我不想在这里纠结太久,但我认为 BLAKE3 验证流(verified streaming)是传输协议的更优替代方案。我们实现了一个名为 iroh-blobs 的协议,它利用 BLAKE3 验证流在 iroh 连接上共享数据块。使用 BLAKE3,我们只需要一个根哈希,既不需要数据块长度参数,也不需要各个数据块的哈希值。这使得 iroh-blobs 能够在无需中间哈希的情况下验证大型数据块的任何部分。关于其工作原理的直观解释,请参阅我关于 BLAKE3 和 Bao 的深度解析,其中涵盖了验证流、外部编码和范围请求。我们仍需努力提高单一大对象多源下载的效率,但这种具备细粒度验证功能的流协议本身,确实优于 BitTorrent 的传输协议。

DHT

DHT(分布式哈希表)

Initially BitTorrent used trackers to get providers for a torrent. What the DHT adds is a way to get providers without having to rely on trackers. This works by computing the SHA-1 hash of the info section, which includes the hashes of all pieces. Then you announce on the DHT that you have content for this hash. The mainline functions for this are announce_peer and get_peers. Each DHT node will store a large set of providers.

最初,BitTorrent 使用追踪器来获取种子文件的提供者。DHT 的加入提供了一种无需依赖追踪器即可获取提供者的方法。其工作原理是计算 info 部分(包含所有数据块的哈希值)的 SHA-1 哈希值,然后在 DHT 上宣告你拥有该哈希对应的内容。Mainline 中用于此目的的函数是 announce_peer 和 get_peers。每个 DHT 节点都会存储大量的提供者信息。

Many more modern protocols also use DHTs for content discovery. But Mainline works extremely well compared to many modern alternatives. It is also an extremely large and stable public DHT deployment, so it is unlikely to go away any time soon. You can look at current mainline statistics. So let’s take a look at why.

许多现代协议也使用 DHT 进行内容发现,但与许多现代替代方案相比,Mainline 的表现极其出色。它也是一个规模巨大且稳定的公共 DHT 部署,因此不太可能在短期内消失。你可以查看当前的 Mainline 统计数据。那么,让我们来看看原因。

Extreme minimalism

极致极简主义

The mainline protocol takes into account that a DHT operates across an extremely large number of nodes. You can’t afford large per-node connection state. Therefore both queries and responses are constrained to fit into single non-fragmented UDP packets. There are a number of other limitations compared to modern DHT implementations that all serve an important purpose: the data stored for a provider is just the public host:port pair of the provider as seen from the DHT node. There is no user defined information in a provider record. For newer extensions like BEP 44 the data is fully self-contained and verifiable, either a tiny piece of data or a signed record, both limited to fit into a typical MTU. You can only store data after first querying the DHT node and returning the short-lived token from its response, proving that at publishing time you can receive packets at your claimed IP address.

Mainline 协议考虑到了 DHT 在海量节点上运行的特性。你无法承担每个节点庞大的连接状态。因此,查询和响应都被限制在单个非分片的 UDP 数据包内。与现代 DHT 实现相比,它还有许多其他限制,但这些限制都有重要的目的:为提供者存储的数据仅仅是 DHT 节点所看到的提供者的公网 host:port 对。提供者记录中没有任何用户自定义信息。对于像 BEP 44 这样的较新扩展,数据是完全自包含且可验证的,要么是一小段数据,要么是一条签名记录,两者都限制在典型的 MTU(最大传输单元)范围内。你必须先查询 DHT 节点并返回其响应中的短期令牌,证明你在发布时能够在声明的 IP 地址接收数据包,之后才能存储数据。

The result of this minimalism is that mainline lookups typically complete in less than a second. Here is a real BEP 44 lookup of a pkarr record using the get_mutable example from n0-mainline.

这种极简主义的结果是,Mainline 的查找通常在不到一秒的时间内完成。以下是使用 n0-mainline 中的 get_mutable 示例对 pkarr 记录进行的真实 BEP 44 查找:

RUST_LOG=debug cargo run --example get_mutable -- \
996a1a05681f92877d5a6ecf9343fba40b8573e5a35ba5773a891ddfb47f9ca6

2026-09-24T06:19:15.626284Z INFO n0_mainline::actor: Mainline DHT started address=0.0.0.0:52517
Looking up mutable item: 996a1a05681f92877d5a6ecf9343fba40b8573e5a35ba5773a891ddfb47f9ca6 ...
Got first result in 37 milliseconds: mutable item: [.. DNS packet bytes omitted ...], seq: 1790095662296563

If you are traumatized by DHT lookups taking forever or timing out: it doesn’t have to be this way. The mainline DHT shows that millisecond lookups are possible at a global scale.

如果你曾因 DHT 查找耗时过长或超时而感到痛苦:其实不必如此。Mainline DHT 表明,在全球范围内实现毫秒级的查找是完全可能的。

Mainline for finding blobs providers

使用 Mainline 寻找数据块提供者

Now that we have established why mainline works so well, let’s see if there is a way to use it for iroh blobs content discovery. Mainline provider records are just an IPv4 host:port pair. But iroh connections are…

既然我们已经确定了 Mainline 为何如此高效,让我们看看是否有一种方法可以将其用于 iroh-blobs 的内容发现。Mainline 的提供者记录仅仅是一个 IPv4 的 host:port 对,但 iroh 连接是……