ParadeDB Optimizes Tantivy to Outperform TIN in Fair BM25 Benchmarks
Key point
ParadeDB optimized Tantivy to close an initial 8x performance gap with PlanetScale's TIN, achieving higher throughput in fair BM25 benchmarks by fixing data locality and algorithm selection.
Details
ParadeDB responded to PlanetScale's benchmark claims regarding its TIN extension by demonstrating that the initial performance gap (where TIN was 8x faster than ParadeDB 0.25) was due to optimization gaps and benchmark anomalies, not architectural differences like ctid vs u32 DocId. Using the StackExchange dataset (150M documents), ParadeDB closed the gap through internal Tantivy optimizations.
Key Optimizations
1. Reduced Random Access in BM25 Scoring
- Problem: Tantivy stored fieldnorms in a separate array, causing scattered lookups in Postgres. In the Hacker News dataset (28.7M docs), fieldnorms accounted for 83% of page accesses.
- Solution: Stored fieldnorm arrays sequentially alongside postings lists in
DocIdorder. - Result: Fieldnorm page accesses dropped from 1,500 to 30. Storage increased by ~9%, but the impact is limited for most terms with short postings lists.
2. Blockmax Pruning Algorithm Selection
- Problem: Multi-term disjunction queries suffered from high CPU overhead with the default Blockmax WAND algorithm.
- Solution: Implemented a heuristic to use MAXSCORE for disjunctions with 3+ terms and dense postings, and WAND otherwise.
- Result: In HN benchmarks (10-term OR query), p50 latency improved ~6x and p95 ~8x. ParadeDB achieved 344.6 QPS vs TIN's 145.7 QPS (2x faster) under fair conditions.
Benchmark Fairness and Accuracy
ParadeDB identified two unintentional anomalies in PlanetScale's original benchmark that favored TIN:
- Syntax Oversight: The original test used ParadeDB's
@@@parser without specifying fields, causing it to search bothidandbodycolumns in StackExchange, while TIN searched only one. ParadeDB switched to native operators (|||,&&&,###) for fair comparison. - Dense-term Elision: TIN skips scoring for terms appearing in >10% of the corpus (
dense_ratio=0.1). This resulted in incorrect rankings: 47.8% of queries had at least one Top 10 result missing from the true Top 10. ParadeDB disabled elision (dense_ratio=2) for accurate BM25 comparison.
StackExchange Mixed Top K Results (150M docs, 8 clients, 5 min):
- ParadeDB: 81.9 QPS
- ParadeDB (stopwords enabled): 161.9 QPS
- TIN (elision disabled,
dense_ratio=2): 35.5 QPS - TIN (elision enabled,
dense_ratio=0.1): 114.4 QPS
Document Identifier Debate
ParadeDB argues that u32 DocId is superior to ctid for columnar integration. While ctid is efficient for simple scoring, it requires additional mapping for columnar operations like Top K by field, range filters, and faceting. u32 DocId maps directly to column row indices, facilitating easier integration with columnar storage.
Release Information
- Version:
0.26.0-rc.2is available for reproduction. - Stable Release: Targeted for next week.
- Action Required: Users must reindex to apply all optimizations, though the update is backwards compatible.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.