AI Briefing
KOSign in

Kapa releases Company Knowledge Bench to evaluate agent retrieval on real-world enterprise data

·2026.10.02 09:00

Key point

The benchmark scores seven retrieval strategies on 1,000 production-derived eval cases, showing that an optimized agentic retriever achieves a score of 0.65 in 5.1 seconds.

Details

Kapa has introduced Company Knowledge Bench, a private benchmark designed to measure how well retrieval systems perform on messy, real-world company knowledge. Unlike public benchmarks that focus on narrow domains or synthetic data, this benchmark uses 1,000 eval cases annotated from actual production queries across developer, sales, and support use cases.

Benchmark Methodology

The benchmark evaluates retrieval based on three core properties:

  • Completeness: The retrieved chunks must fully answer the query.
  • Minimality: No extra chunks should be included, as they increase cost and reasoning difficulty.
  • Source Preference: The system must prioritize authoritative and current sources (e.g., a dedicated reference page over a community forum post).

To create the benchmark at scale, Kapa used a two-stage agent process validated against 170 human-labeled eval cases. Candidate agents scan for relevant chunks, and criteria agents generate the final retrieval criteria based on a strict handbook.

Retrieval Performance Results

Kapa tested seven different retrieval strategies, ranging from traditional RAG pipelines to agentic approaches. The results highlight significant trade-offs between accuracy, latency, and cost:

  • Hybrid search (Baseline): Scored 0.41 with the lowest latency (0.4 s) but poor precision.
  • Hybrid search + rerank: Adding a reranker (rerank-2) improved the score to 0.50 with minimal latency increase (0.7 s).
  • Query decomposition: Using gpt-6.1-luna to split queries improved the score to 0.56 but doubled latency to 1.8 s.
  • Agentic retrievers: gpt-6.1-sol with grep tools achieved a score of 0.61 but with high latency (17 s) and low precision (4%). gpt-6.1-luna scored 0.54.
  • Kapa Default: A fixed pipeline achieved a score of 0.61 in 3.3 s with 26% precision.
  • Kapa Deep: An optimized agentic retriever achieved the highest score of 0.65 in ~5 s with the highest precision (26%).

Cost and Efficiency

The study highlights that agentic retrievers can be costly due to token volume. The gpt-6.1-sol grep agent returns ~40,000 tokens per query, costing ~$77.58 per 1,000 queries in input tokens alone (assuming $2/M input). In contrast, Kapa Deep returns fewer tokens, costing ~$10.20 per 1,000 queries. The benchmark concludes that optimized agentic retrieval offers the best balance of accuracy, latency, and cost.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.