Release notes
What changed in each tatami release.
The authoritative, commit-level history lives in CHANGELOG.md and on the releases page. This page summarises each version.
v0.2.0
The serving release. v0.1.0 proved the format and a single-process search index; v0.2.0 turns it into something a fleet can run, with a routed broker over a fan of cold shards, an aggregator that merges an exact fleet-wide top-k across many brokers, search-only segments that drop the body they never serve, and an HTTP server that answers thousands of concurrent queries with the latency budget intact and the memory bounded.
- A routed, pruned broker. A
Clusterserves a large fan of cold shards behind one query, opening only a bounded working set at a time and visiting only the shards that can contribute. A routing index maps each term to the shards that hold it with a per-shard impact bound, so a query walks shards in descending bound order and stops the moment the next bound cannot beat the current k-th best. Because the ranking bound is a true upper bound, the pruned result is byte-identical to a full fan-out. On a real shard split into 254 shards with a 128-segment cache, a selective keyword query prunes to nine or ten candidate shards and keyword retrieval runs at a p99 of about 1.5 milliseconds. - Exact cross-shard ranking. Global statistics injection scores every shard against the same corpus-wide document count and per-term document frequency, so the merged top-k equals the top-k of a single index over every shard. This closes the best-effort cross-segment merge v0.1.0 left open.
- Search-only segments. A search segment can drop the document body it only ever needed to build the postings, keeping a short precomputed snippet in its place. Retrieval is byte-identical to a full-document segment because the body is still tokenised before it is dropped. On the production ccrawl shard the search-only segment is 42.8 percent smaller than the full-document one.
- An aggregator tier. An
Aggregatorfans a query out to many leaf brokers concurrently, scores every leaf against fleet-wide statistics, and merges the leaves' partial top-k lists into one fleet-wide top-k that dedups a re-crawled page by stable id and is exact. A single-root merge over a hundred thousand shards' worth of leaves is the only part that grows with fleet size, and a tree of aggregators clears it, projecting a fleet p99 of about 1.3 milliseconds at a hundred thousand shards. - The serve command.
tatami serve <dir>runs an HTTP server over a directory of search segments:GET /search?q=&k=returns ranked JSON,/healthzis a liveness probe, and/statsreports the broker shape and the serving counters. The broker answers each query without a shared lock, so one process handles many concurrent requests; a smart segment cache keeps only the working set resident, admission control sheds a burst past the cap with a 503 rather than queuing without bound, and a per-request deadline bounds a stalled query with a 504. On a real shard, single-keyword serving holds a p99 of about 1.5 milliseconds at over 31,000 queries per second, the resident memory stays bounded by the cache cap through thousands of concurrent queries, and a concurrent answer is identical to the single-threaded one.
v0.1.0
The first release. tatami is a single-file columnar storage format that stores a crawl corpus compactly and doubles the same file as a keyword search index. It ships the full container, the encoding and compression stack, the indexing and pruning structures, the collection layer, the Parquet bridge, and the search-segment role, proven on real Common Crawl data.
- The container. A fixed 64-byte header, row groups of column chunks, an optional blob, dictionary, and index region, and a self-describing footer written last so a reader learns the whole layout from one tail read. The magic
TAT1sits at both ends, every page header is uncompressed so a reader can stride over pages without decoding, and a CRC32C guards every page and the footer. - An encoding cascade. Each page is encoded with the cheapest scheme that fits (bit-packing with a frame of reference, delta, run-length, dictionary, group-varint, PForDelta, FSST for strings, bitmaps) and then compressed with zstd. The choice is per page and deterministic, so the same input produces a byte-identical file.
- Blob separation and shared dictionaries. Large payloads like markdown bodies are written to their own region and referenced by offset, so a metadata scan never reads a body. A trained dictionary spans every row group rather than being rebuilt per group, which is where the format pulls ahead of a per-group layout.
- Indexing and pruning. Zone maps on every chunk, opt-in bloom filters, and a sparse primary-key index on a sorted file. A scan pushes a predicate down to the group and page level; a lookup is a bounded seek that reads one page index and one data page regardless of file size.
- Collections. A
tatami.manifestcatalogs a directory of files into one logical dataset, with a rollup of each file's key range and zone statistics so a query prunes whole files before opening them. Add, list, compact from the CLI; scan, look up, and merge from the Go API. - The Parquet bridge.
tatami convertre-encodes an existing Parquet crawl shard as tatami without a producer change. On a real Common Crawl markdown shard the file comes out about 26 percent smaller than the zstd Parquet source, with every body preserved byte for byte. - The search-segment role. A header bit turns a file into a search index: an inverted region over the forward columns, a posting codec with block-max WAND retrieval, BM25 ranking with an exact top-k, and a forward fetch that reads a hit's url and title with two cached column reads. On a real shard, 20246 documents and 1.4 million terms, keyword queries return with a p99 of 237 microseconds.
- Tiered merge and serving at scale. Deletions clear a bit in a live-docs bitset and are honored at query time without rewriting a sealed file. A tiered merge folds small segments into large ones and drops the deleted documents. An
Indexserves many segments behind one query with a global top-k and stable-id dedup; split across twenty segments, fan-out keyword retrieval stays at a p99 of 465 microseconds. - Packaged everywhere. Archives for Linux, macOS, Windows, and FreeBSD on amd64 and arm64,
.deb/.rpm/.apkpackages, a multi-arch GHCR image, Homebrew and Scoop entries, checksums, SBOMs, and a cosign signature.