AI Briefing
KOSign in

Yeogi-eottae Builds Agentic RAG Operations and Evaluation System Using Langfuse and FastMCP

·2026.10.10 09:01

Key point

Accuracy improved by 3–9%p through LLM-as-a-judge and MCP specification optimization, applied to three internal services.

1 / 6

Details

The Yeogi-eottae Common Platform Development Team shared a case study on adopting Langfuse and FastMCP to ensure observability and operational stability for their internal knowledge-based Agentic RAG system. The focus was on building a pipeline that goes beyond simple search functionality to verify answer accuracy and guide agents to use tools effectively.

Building LLM Observability and Evaluation Pipeline

LLM and RAG systems are difficult to evaluate due to the nature of natural language outputs, and cost tracking is often opaque. To address this, the open-source observability platform Langfuse was introduced to record and monitor prompts, responses, token usage, and costs at the trace level. The evaluation was designed as a dual structure combining retrieval metrics (Recall@5, Hit Rate, MRR) and LLM-as-a-judge. Retrieval metrics verify the retriever's document recall capability, while LLM-as-a-judge scores the accuracy of generated answers on a scale from 0 to 1. Specifically, the judge model was provided with a ground truth answer key to score against, enhancing score stability.

FastMCP Serving and Tool Orchestration

FastMCP was used to provide search tools to the agent, adopting the HTTP Streamable transport method to centralize permission management on the server side. Server instructions were delivered during the MCP Protocol initialization phase to ensure the agent followed the intended search flow (semantic search → relationship exploration → content reading). This prevented the issue of agents ignoring tool call sequences and providing immediate answers.

Performance Improvement Through Specification Optimization

Accuracy was improved by optimizing only the MCP specification (instructions, docstrings, parameters) without changing the search logic. Lighter models are more prone to the 'lost in the middle' phenomenon, where instructions in the middle of long prompts are missed; therefore, redundant instructions were removed and prompts were kept concise. Additionally, the method of receiving multiple documents at once was changed to process them one by one, and path-based identifiers were replaced with unique IDs to reduce model interpretation errors. These specification refinements alone increased accuracy by 3–9%p and reduced unnecessary round-trip calls, thereby lowering token usage.

Internal Service Application Status

The built Atlas Retriever MCP is currently in operation across three internal services. It is used for generating answer drafts for the Slack-based technical support bot, as a knowledge search tool for the internal AI agent platform YAPP, and for integration with coding tools in personal development environments. It targets enterprise-wide Confluence and Jira documents, reflecting the latest knowledge through incremental updates three times a day.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.