Exhibition Development Team Enhances Incident Response Efficiency with Grafana-Based Observability Board
Key point
Sharing a case study on improving internal observability and fostering contributions by leveraging existing infrastructure and AI agents.
Details
The Exhibition Development Team serves as a core service handling Yeogi Eottae's key traffic, where large-scale traffic processing and stable communication with linked systems are critical. After a team transition, they built an observability board based on Grafana and Grafana Tempo/Loki to address the limitations of the existing Pinpoint APM while leveraging existing infrastructure.
The observability board was designed to instantly identify the core question of incident response: 'Is it our service's problem, or the linked service's problem?' It intuitively displays request/success/failure TPS as numbers, separates Incoming and Outgoing traffic to allow immediate isolation of average latency and errors, and provides a one-click navigation feature from error metrics to actual logs and traces, eliminating context switching.
During implementation, Claude Code and the internal AI agent Yapp MCP were used to automate repetitive tasks such as metric exploration and query adjustment, allowing engineers to focus on core design.
The shared board led to additional contributions from team members, resulting in the creation of additional boards for identifying anomalies across microservices and for event response. However, due to the inability to set measurement criteria before the work, quantitative comparison of effects was difficult. The team left the task of advancing the AI agent to perform RCA (Root Cause Analysis) as a future goal.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.