We used a database as a message queue. Now we use Kafka

We used a database as a message queue. Now we use Kafka

我们曾用数据库做消息队列,现在我们改用 Kafka

FoundationDB is the only database we use. This should surprise you since FoundationDB is pretty barebones, just a key-value store. It stores everything for us: tenants, object metadata, the replication log for data distributed across regions, etc. We also use it as a queue to handle async tasks, à la QuiCK, the queuing system Apple uses for CloudKit. This has scaled very nicely. I’m not surprised; it’s the same tech behind iCloud, a platform with at least 900 million users. Furthermore, keeping the queue inside FoundationDB means all transactions stay in the database, eliminating the dual-write problem.

FoundationDB 是我们使用的唯一数据库。这可能会让你感到惊讶,因为 FoundationDB 非常精简,仅仅是一个键值存储。它为我们存储了一切:租户、对象元数据、跨区域数据分发的复制日志等。我们还将其用作处理异步任务的队列,类似于苹果用于 CloudKit 的队列系统 QuiCK。它的扩展性非常好。我对此并不感到惊讶;它正是 iCloud 背后的技术,而 iCloud 是一个拥有至少 9 亿用户的平台。此外,将队列保留在 FoundationDB 中意味着所有事务都留在数据库内,从而消除了双写问题。

So why start using Kafka now? We’ve seen a few issues using our database as a message queue: Scheduling requires many writes and scans, which puts read load on FoundationDB that directly competes with user requests. Each task is expensive and needs multiple writes to complete (enqueue, claim, lease, etc). We have ever more tasks as we add more features. New team members have to learn all the custom code resulting from actually implementing the QuiCK paper. There’s no standard implementation, even though it’s a well known pattern in theory. Finesse is not something you can learn from a paper.

那么,为什么现在要开始使用 Kafka 呢?在使用数据库作为消息队列时,我们遇到了一些问题:调度需要大量的写入和扫描,这给 FoundationDB 带来了直接与用户请求竞争的读取负载。每个任务都很昂贵,需要多次写入才能完成(入队、认领、租约等)。随着我们增加更多功能,任务也越来越多。新团队成员必须学习实现 QuiCK 论文所带来的所有自定义代码。尽管这在理论上是一个众所周知的模式,但并没有标准实现。技巧是无法从论文中学到的。

This isn’t a story of a neat 1:1 replacement. We still have the queues in FoundationDB. We moved asynchronous tasks like garbage collection to Kafka, we can reduce the read and write load on FDB and shave off a good amount of that pesky custom code. Read more to see how it all turned out for us!

这不是一个简单的 1:1 替换故事。我们仍然在 FoundationDB 中保留了队列。我们将垃圾回收等异步任务迁移到了 Kafka,这样可以减少 FDB 的读写负载,并剔除掉大量恼人的自定义代码。继续阅读以了解我们的最终成果!


note 注

You might tell us we should have “just used Kafka” the whole time. Beyond the fact that you used the j-word: have you ever waited for your not-even-that-big broker to catch up on a cold start? Do you know what a zookeeper is and why you don’t pay to take care of the animals? The poor zookeeper can’t even pet them. Have you ever felt like a plastic bag drifting through the wind but unable to start again because of the sheer madness that comes with spending months permuting JVM flags to try to eke out a spectre’s worth of performance so that your servers aren’t constantly on fire? No? Just me?

你可能会告诉我们,我们应该“一直使用 Kafka”。除了你用了那个“j”开头的词(指 Java)之外:你有没有经历过等待一个规模并不大的代理(broker)在冷启动后追赶数据?你知道什么是 Zookeeper,以及为什么你不用付钱去照顾那些动物吗?可怜的管理员甚至不能抚摸它们。你有没有感觉自己像随风飘荡的塑料袋,却因为花费数月调整 JVM 参数以榨取那一点点性能,只为让服务器不至于持续过载而陷入疯狂,以至于无法重新开始?没有?只有我吗?

Either way we kinda wanted to avoid Kafka because running it yourself is the administrative experience of finding yourself turned into a monstrous vermin and everyone around you is mildly annoyed at your experience and asking you to move on with life instead of understanding that you can’t work anymore because your arms have turned into dozens of legs. By the way, that’s actually what people mean when they call something “kafkaesque”, not a Qu’vatlh of paperwork.

无论如何,我们确实想避免使用 Kafka,因为自行运维它的管理体验就像是你发现自己变成了一只巨大的害虫,周围的人对你的遭遇感到轻微的厌烦,并要求你继续生活,而不是理解你因为手臂变成了几十条腿而无法再工作。顺便说一句,这才是人们称某事为“卡夫卡式”(kafkaesque)的真正含义,而不是指堆积如山的文书工作。


Beyond the naïve way to build a queue on FoundationDB

超越在 FoundationDB 上构建队列的“天真”方式

So you need a queue. The FoundationDB docs contain tutorials for making simple queues in multiple languages. At a high level you put messages in on one end of the keyspace (namespace for keys) and then read them out of the other end of the keyspace. This works fairly well (if you’ve ever used Sidekiq, this model should be very familiar), but the main problems come with naming the entries in the queue. The naïve way to do it is to use the FoundationDB equivalent of MySQL’s AUTO INCREMENT where you assign each queue item its own atomically increasing integer ID, but what happens when you have more than one producer?

你需要一个队列。FoundationDB 文档中包含了用多种语言制作简单队列的教程。从宏观上看,你将消息放入键空间(键的命名空间)的一端,然后从另一端读取出来。这运行得相当不错(如果你用过 Sidekiq,这个模型应该很熟悉),但主要问题在于如何命名队列中的条目。最天真的做法是使用 FoundationDB 中等同于 MySQL 的 AUTO INCREMENT 的功能,为每个队列项分配一个原子递增的整数 ID,但当你有多个生产者时会发生什么呢?

Given a sufficiently distributed system, it’s easy for two jobs to have conflicting IDs, such as two tasks getting the ID 67 and conflicting with eachother on insert. Sure, with enough work you can random or UUID your way out of this, but the core problem is that consuming work deletes it from the database. If a worker dies while it’s processing an item, there’s no way for another worker to retry. Once the worker consumes a job, it’s no longer in the queue and that job dies with it.

在足够分布式的系统中,两个作业很容易产生冲突的 ID,例如两个任务都获得了 ID 67,并在插入时发生冲突。当然,通过足够的工作,你可以使用随机数或 UUID 来解决这个问题,但核心问题在于,消费作业会将其从数据库中删除。如果一个工作进程在处理某个项目时死亡,其他工作进程将无法重试。一旦工作进程消费了一个作业,它就不再在队列中,该作业也随之消失。


We implemented an Apple paper that describes how to implement durable message queues on top of FoundationDB

我们实现了一篇苹果论文,描述了如何在 FoundationDB 之上实现持久化消息队列

Turns out Apple has thought about this and implemented it for CloudKit with a system they call QuiCK, or a Queueing System in CloudKit. Our industry has silly names for things. QuiCK is a robust queuing system that uses FoundationDB’s Record Layer to implement a message queue based on time. To understand why this is such a galaxy-brained genius move, let’s take a sidestep into how time works in distributed systems.

事实证明,苹果已经考虑过这个问题,并为 CloudKit 实现了一个名为 QuiCK(CloudKit 中的队列系统)的系统。我们这个行业对事物的命名总是很滑稽。QuiCK 是一个强大的队列系统,它利用 FoundationDB 的 Record Layer 来实现基于时间的队列。为了理解为什么这是一个天才般的举动,让我们先侧面了解一下时间在分布式系统中是如何运作的。