Musinsa Resolves Service Outage Caused by Logging: Securing Asynchronous Logging and trace_id with 30 Lines of Code
Key point
Resolved a connection pool exhaustion outage caused by stdout buffer saturation using a custom Appender, achieving both asynchronous logging and trace_id preservation.
Details
The Musinsa PEL organization experienced an outage in a high-traffic service where request threads were blocked due to stdout buffer saturation, leading to connection pool exhaustion. This was a side effect of the company-wide standard guide that prohibited the use of AsyncAppender for observability purposes and enforced synchronous logging.
Cause of the Outage: Limitations of Synchronous Logging
Synchronous logging requires the request thread to write directly to stdout. When log volume surged to 100 times the normal level, the 64KB pipe buffer filled up, causing threads to wait until the 30-second connection acquisition timeout. This resulted in reduced availability across the entire service.
Solution: TraceAwareAsyncAppender
The development team introduced a custom TraceAwareAsyncAppender that extends AsyncAppender. This approach captures the trace_id and span_id from OpenTelemetry in the caller thread before queuing the log event. With approximately 30 lines of code, it simultaneously secured the performance benefits of asynchronous logging and the observability of distributed tracing.
Verification Results
In load tests, even under extreme log volumes of 15,000 lines per second, the trace_id missing rate in the output logs was 0%. Thanks to the neverBlock policy, request threads did not wait even if some logs were dropped, and the p95 response time remained at 29ms.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.