RIP, vector database
RIP, vector database
再见,向量数据库
September 30, 2026 • Dan Harrison (Engineer) 2026年9月30日 • Dan Harrison(工程师)
We are changing turbopuffer’s storage architecture to take search to the next level. turbopuffer v3 changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to make search faster in every respect — including text, regex, and vector search — but it also lays the foundation to move many more SQL queries to turbopuffer and make them fast. 我们正在改变 turbopuffer 的存储架构,以将搜索能力提升到一个新的水平。turbopuffer v3 改变了文档和索引在 turbopuffer 中的布局、写入、压缩和查询方式。这将使我们在各个方面(包括文本、正则表达式和向量搜索)都能实现更快的搜索,同时也为将更多 SQL 查询迁移到 turbopuffer 并实现高性能奠定了基础。
turbopuffer launched as a serverless vector database (v1), highly specialized to the task of serving extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion. turbopuffer 最初作为无服务器向量数据库(v1)发布,专注于提供极其廉价且速度尚可的向量搜索任务。以对象存储作为事实来源(source of truth)保证了经济性,而分层 NVMe SSD/内存缓存则保证了性能。这些特定权衡的价值已得到我们最早的客户(包括 Cursor 和 Notion)的验证。
turbopuffer evolved to have very strong text and regex search (v2), and is being used for many non-search use cases, like Linear’s syncing engine. The query engine has evolved along the way to support all of these query plans, but the storage architecture has remained largely unchanged: the ANN vector index was and still is the primary index around which all other indexes and query plans revolve. This design has constrained several query plans, like GROUP BY and aggregations. turbopuffer 后来演进出了非常强大的文本和正则表达式搜索功能(v2),并被用于许多非搜索场景,例如 Linear 的同步引擎。查询引擎在此过程中不断演进以支持所有这些查询计划,但存储架构却基本保持不变:ANN 向量索引过去是、现在仍然是所有其他索引和查询计划围绕的核心主索引。这种设计限制了诸如 GROUP BY 和聚合等多种查询计划。
We’ve pushed the vector-primary architecture as far as we can, and it’s time to move on. We’re in the process of moving to a new primary index, and making ANN “just another” secondary index. We thought it might be fun to open up the doors and let you follow along. For this first update, we’ll set the stage with why we’re doing this in the first place. Walk with me on a short journey from tpuf v1 to today. 我们已经将“向量优先”架构推向了极限,现在是时候做出改变了。我们正在转向一种新的主索引,并将 ANN 降级为“仅仅是”另一个二级索引。我们认为,向大家公开这一过程并邀请你们共同见证会很有趣。在第一次更新中,我们将先说明为什么要这样做。请跟随我,回顾一下从 tpuf v1 到今天的简短历程。
v1: an ID and a vector
v1:一个 ID 和一个向量
In the first version of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing wisdom at the time was graph-based vector indexes, but a hierarchical clustering index plays better with object storage. We started with SPANN, and eventually migrated to SPFresh to support incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree with a single root. 在 turbopuffer 的第一个版本中,文档仅由一个 ID 和一个向量组成。当时的主流观点是基于图的向量索引,但分层聚类索引在对象存储上表现更好。我们从 SPANN 开始,最终迁移到 SPFresh 以支持增量索引。向量被聚类成组,组的质心又被进一步聚类,如此重复,形成了一棵具有单一根节点的树。
We implemented this on top of a storage layer presenting as a key-value map, with sorted and unique keys. Each cluster is given a ClusterId, and vectors within each cluster are given a dense LocalId. 我们在一个表现为键值映射(key-value map)的存储层之上实现了这一点,该层具有排序且唯一的键。每个集群被分配一个 ClusterId,集群内的向量则被分配一个密集的 LocalId。
As you can see above, everything is keyed by ClusterId and LocalId (e.g. C0L1), which together we call the ANN address. This is what we mean when we say the ANN index is the primary index. 如上所示,所有内容都以 ClusterId 和 LocalId(例如 C0L1)为键,我们将其统称为 ANN 地址。这就是我们所说的“ANN 索引是主索引”的含义。
v2: attribute filtering and full-text search
v2:属性过滤与全文搜索
Two new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search. 两个新的查询计划标志着 turbopuffer 从 v1 到 v2 的非正式过渡:属性过滤和全文搜索。
Attribute filtering 属性过滤
Naturally, customers wanted to be able to add attribute values and filter vector searches on them. To make filtering fast and high-recall, we modeled these as an inverted index that maps an attribute value to the ANN address of the documents that contain it. 自然地,客户希望能够添加属性值并据此过滤向量搜索。为了使过滤快速且保持高召回率,我们将这些属性建模为倒排索引,将属性值映射到包含该属性的文档的 ANN 地址。
For projections (include_attributes), we also stored the document attributes alongside the ID and the vector. 对于投影(include_attributes),我们还将文档属性与 ID 和向量一起存储。
Full-text search 全文搜索
BM25 full-text search was another obvious and much-demanded query plan. Similar to attribute search, full-text search works by first finding the documents that have the query term present (commonly called “postings”). For an FTS index, we also include the (term count, document length) metadata necessary for BM25 scoring. BM25 全文搜索是另一个显而易见且需求量很大的查询计划。与属性搜索类似,全文搜索的工作原理是首先找到包含查询词的文档(通常称为“倒排表”)。对于 FTS 索引,我们还包含了 BM25 评分所需的(词频、文档长度)元数据。
Over time, we’ve shipped several other index structures and query engines: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built around the same vector-primary storage layout. 随着时间的推移,我们发布了其他几种索引结构和查询引擎:聚合、正则表达式搜索、模糊匹配、稀疏向量搜索和属性排序——所有这些都是围绕相同的“向量优先”存储布局构建的。
The problem with a vector primary index
“向量优先”主索引的问题
The ANN primary index has largely remained intact until today for one simple reason: it works really, really well for ANN search on object storage. On top of this architecture, we’ve pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. Any significant change here risks introducing regressions in ANN performance. ANN 主索引至今基本保持不变,原因很简单:它在对象存储上的 ANN 搜索表现确实非常出色。基于这种架构,我们将向量搜索扩展到了拥有超过 1000 亿个向量的单一索引,实现了 200 毫秒的 p99 读取延迟和每秒 1000 次以上的查询吞吐量。在此架构上进行任何重大更改都有可能导致 ANN 性能下降。
However, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization. 然而,这种布局在三个主要方面阻碍了我们在所支持的非向量查询类型上达到顶尖水平:存储放大、写入放大和有限的向量化。
Storage amplification 存储放大