Leviathan Releases Local Log Indexer for Agents
Key point
Achieved 99% reduction in context tokens compared to grep on 1 million records with 98.5% hit@1 accuracy.
Details
The Rust-based single binary tool Leviathan has been released. It converts large-scale data such as JSONL, CSV, and SQLite into SQLite FTS5 indexes, delivering only compressed search results of a few hundred tokens to AI agents using BM25 lexical search without embeddings.
Performance and Efficiency
Benchmark results on 1 million synthetic records released by the developer show that Leviathan dramatically reduces agent context usage compared to grep-based approaches.
- Token Efficiency: On 1 million records, using full grep history consumes 209,412 tokens, whereas Leviathan uses 436 tokens, achieving approximately 99% reduction.
- Search Accuracy: Recorded hit@1 98.5% and hit@5 99.0%, showing higher accuracy than grep-based methods (96.0%, 83.0%).
- Latency: The p50 latency for searching 1 million records is 33.4ms.
Features and Limitations
- Lightweight: Runs locally without network or DB connection information and has no runtime dependencies. Index size is approximately 1.8x the original.
- Agent Integration: Can be integrated via Claude Skills or MCP servers (
leviathan mcp) with 4 read-only tools such assearchandget. - Korean Support: The default
porter unicode61tokenizer does not support Korean morphological analysis, limiting particle processing. Experimental auxiliary field generation using a morphological analyzer is possible. - Data Characteristics: Optimized for structures with many records for specific targets (e.g., equipment maintenance records) and is currently in the early v0.1.0 stage.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.