What Happened When We Loaded Madrid’s GTFS Data Into a Reactive Knowledge Graph

What Happened When We Loaded Madrid’s GTFS Data Into a Reactive Knowledge Graph

当我们将马德里的 GTFS 数据加载到响应式知识图谱中时发生了什么

I’ve been building .me around a fairly simple idea: Knowledge is not just data. It is data plus the relationships that make changes meaningful. Most of my previous tests of this idea used synthetic workloads. That’s useful when testing things like fan-out, dependency depth, recomputation, and early cutoff. But synthetic tests eventually raise an uncomfortable question: What happens with somebody else’s data? So I took the GTFS Madrid workload used in research on incremental knowledge graph construction and adapted it to run on .me. What started as a benchmark ended up being more interesting than I expected. We effectively built a small, evolving knowledge graph of Madrid’s transportation data inside .me. And then we changed the world underneath it.

我一直围绕一个相当简单的想法构建 .me:知识不仅仅是数据。它是数据加上使变化具有意义的关系。我之前对这个想法的大多数测试都使用了合成工作负载。这在测试扇出(fan-out)、依赖深度、重新计算和提前截止等功能时很有用。但合成测试最终会引出一个令人不安的问题:如果是别人的数据会怎样?因此,我采用了在增量知识图谱构建研究中使用的 GTFS 马德里工作负载,并对其进行了适配,使其能够在 .me 上运行。最初只是一个基准测试,结果却比我预期的更有趣。我们实际上在 .me 内部构建了一个小型、不断演进的马德里交通数据知识图谱。然后,我们改变了它底层的世界。

⸻

First: what did .me actually “learn”? “Learn” is probably the wrong technical word. There was no machine learning involved. No model was trained. Instead, .me acquired a structured representation of a domain. GTFS contains transportation data describing things such as: routes, trips, services, calendars, stop times. The adapter loaded those entities into the .me namespace.

首先:.me 到底“学习”到了什么?“学习”这个词在技术上可能用得不对。这里不涉及任何机器学习,也没有训练模型。相反,.me 获取了一个领域的结构化表示。GTFS 包含描述以下内容的交通数据:路线、行程、服务、日历、停靠时间。适配器将这些实体加载到了 .me 的命名空间中。

Conceptually: 从概念上讲:

gtfs
├── services
│ ├── service_1
│ ├── service_2
│ └── ...
├── routes
│ ├── route_1
│ └── ...
├── trips
│ ├── trip_1
│ ├── trip_2
│ └── ...
└── stopTimes
    ├── ...
    └── ...

But simply storing GTFS records in a tree wouldn’t be particularly interesting. The important part was connecting them.

但仅仅将 GTFS 记录存储在树中并不会特别有趣。重要的部分是将它们连接起来。

⸻

From data to dependencies 从数据到依赖关系

A transportation system contains relationships. A trip belongs to a route. A trip uses a service. A service calendar determines when that trip can operate. Stop times belong to trips. Conceptually: service → trip → stop_time. The adapter expressed relevant relationships as explicit .me dependencies. A tiny version of the same idea looks like this: me.A(2), me.B(3), me.C = A + B. C is no longer merely another stored value. It is knowledge derived from A and B. If we change me.A(5), .me knows that C depends on A. We can inspect that relationship: me.explain("C") and get information corresponding to: sourcePath: A, dependsOn: [A, B], k: 1, recomputed: [C]. The GTFS adapter takes this basic mechanism and applies it across a much larger domain.

交通系统包含各种关系。行程属于路线,行程使用服务,服务日历决定了行程何时运行,停靠时间属于行程。从概念上讲:服务 → 行程 → 停靠时间。适配器将相关关系表达为显式的 .me 依赖项。这个想法的一个微小版本如下所示:me.A(2), me.B(3), me.C = A + B。C 不再仅仅是另一个存储的值,它是从 A 和 B 派生出来的知识。如果我们更改 me.A(5),.me 知道 C 依赖于 A。我们可以检查这种关系:me.explain("C") 并获得相应的信息:sourcePath: A, dependsOn: [A, B], k: 1, recomputed: [C]。GTFS 适配器采用了这种基本机制,并将其应用于更大的领域。

⸻

Building the Madrid knowledge universe 构建马德里知识宇宙

I ran the benchmark at three scales. At scale 1, the loaded GTFS snapshot contained 2,364 stop times. At scale 10: 23,640 stop times. At scale 100: 236,400 stop times. The scale-100 base snapshot included approximately: routes 1,300, trips 13,000, stop_times 236,400, calendar 500, calendar_dates 7,000. Alongside that source knowledge, the adapter created indexes and derived relationships used by the dependency graph. So at this point .me wasn’t looking at a toy A → B → C example anymore. It had a structured transportation universe containing hundreds of thousands of GTFS records and relationships.

我在三个规模上运行了基准测试。在规模 1 下,加载的 GTFS 快照包含 2,364 个停靠时间。在规模 10 下:23,640 个停靠时间。在规模 100 下:236,400 个停靠时间。规模 100 的基础快照大约包括:路线 1,300 条,行程 13,000 个,停靠时间 236,400 个,日历 500 个,日历日期 7,000 个。除了这些源知识外,适配器还创建了依赖图使用的索引和派生关系。因此,此时的 .me 不再是看着一个简单的 A → B → C 示例,它拥有一个包含数十万条 GTFS 记录和关系的结构化交通宇宙。

⸻

Then we changed Madrid 然后我们改变了马德里

This is where the experiment becomes interesting. The GTFS benchmark doesn’t only provide a static dataset. It provides evolving versions of that dataset. I compared the base version with seed0 and converted the differences into neutral mutations: CREATE, UPDATE, DELETE. Then those mutations were applied to .me one by one. Across the three scales, the complete sweeps contained: Scale 1: 151 mutations; Scale 10: 1,304 mutations; Scale 100: 13,401 mutations. At scale 100: 13,401 / 13,401 mutations were measurable. No mutations were skipped. Now we could ask a much more interesting question than “how fast can you read a node?” We could ask: When one fact in this knowledge universe changes, how much other knowledge does that change actually affect?

这就是实验变得有趣的地方。GTFS 基准测试不仅提供静态数据集,还提供该数据集的演进版本。我将基础版本与 seed0 进行了比较,并将差异转换为中性突变:创建 (CREATE)、更新 (UPDATE)、删除 (DELETE)。然后,这些突变被逐一应用于 .me。在三个规模上,完整的扫描包含:规模 1:151 个突变;规模 10:1,304 个突变;规模 100:13,401 个突变。在规模 100 下:13,401 / 13,401 个突变均可测量,没有突变被跳过。现在我们可以问一个比“读取节点有多快?”更有趣的问题:当这个知识宇宙中的一个事实发生变化时,这种变化实际上影响了多少其他知识?

⸻

n versus k n 与 k

This distinction is central to .me. Let: n = total knowledge loaded into the system, and k = knowledge affected by a particular mutation. These are not necessarily the same variable. Imagine a graph containing a million pieces of knowledge. Changing one value doesn’t necessarily affect a million things. Maybe it affects three. Maybe twenty. Maybe 100,000. The topology determines that. So instead of assuming that a mutation should somehow involve the entire knowledge universe, .me keeps explicit dependencies and follows the affected region.

这种区别对 .me 至关重要。设:n = 加载到系统中的总知识量,k = 特定突变影响的知识量。这些变量不一定相同。想象一个包含一百万条知识的图谱。改变一个值并不一定会影响一百万件事。也许它影响三个,也许二十个,也许十万个。拓扑结构决定了这一点。因此,.me 不会假设突变必须涉及整个知识宇宙,而是保留显式的依赖关系并跟踪受影响的区域。

⸻

The surprising part 令人惊讶的部分

Here are the results: 以下是结果:

ScaleStop TimesMutationsk p50k p95k maxRecompute p50Recompute p95
12,364151118.5270.004 ms0.073 ms
1023,6401,304118270.003 ms0.086 ms
100236,40013,401118270.003 ms0.131 ms

Put differently: 换句话说: n: 2,364 → 23,640 → 236,400 k p95: 18.5 → 18 → 18 k max: 27 → 27 → 27

The dataset became approximately 100× larger. The affected-set distribution essentially didn’t. That is the result I find most interesting.

数据集大约变大了 100 倍,但受影响集合的分布基本没有变化。这是我觉得最有趣的结果。

⸻

DELETE was even more stable DELETE 操作更加稳定

DELETE mutations gave us a particularly clean pattern. Across scale 1, scale 10, and scale 100: DELETE k p50 = 18, DELETE k max = 19. Exactly the same. So while the surrounding knowledge universe became roughly two orders of magnitude larger, the dependency neighborhood touched by those mutations stayed local.

DELETE 突变给了我们一个特别清晰的模式。在规模 1、规模 10 和规模 100 中:DELETE k p50 = 18,DELETE k max = 19。完全相同。因此,尽管周围的知识宇宙大约扩大了两个数量级,但这些突变所触及的依赖邻域仍然保持在局部。

⸻

This does NOT mean everything is O(k) 这并不意味着一切都是 O(k)

This is where I want to be careful. It would be tempting to look at these numbers and say: “.me mutations are O(k).” We haven’t demonstrated that. In fact, the benchmark exposed a bottleneck that makes that claim inappropriate. There are at least two different things happening when we mutate this knowledge universe: mutation = structural application + dependency propagation. Dependency recomputation remained very small. Structural mutation handling did not always behave that way. At scale 100, DELETE apply operations were often measured in seconds.

在这里我需要谨慎。人们很容易看到这些数字并说:“.me 的突变是 O(k) 的。”我们还没有证明这一点。事实上,基准测试暴露了一个瓶颈,使得这种说法不恰当。当我们改变这个知识宇宙时,至少发生了两件不同的事情:突变 = 结构应用 + 依赖传播。依赖重新计算仍然非常小,但结构突变处理并不总是表现得那样。在规模 100 下,DELETE 应用操作通常以秒为单位进行测量。