How Does Oliveyoung Manage Incidents?
Key point
Oliveyoung operates a systematic process from incident policy establishment to post-incident management to ensure rapid response to outages and prevent recurrence.
Details
Oliveyoung has established and operates an Incident policy to minimize unexpected service malfunctions and revenue loss. It identifies the Usecase of all systems and defines a CSP (Critical Serving Path) according to importance, focusing on reducing the fatigue of incident response.
Incident levels are defined based on revenue and customer impact, and are determined through consultation with relevant departments (MD, Sales, Marketing, SCM, etc.). When an incident occurs, a channel is created via Slack, and immediate notifications are propagated to relevant personnel through an AWS Lambda-based automation system.
After an incident is resolved, the process does not simply end there; the following post-incident activities are carried out.
- Writing an Incident Report: The 5 Why Questions method is used to derive the Root Cause.
- Incident Review Meeting: A monthly meeting is held to share causes and discuss measures to prevent recurrence. The culture here aims to focus on finding solutions rather than assigning blame.
- Managing Recurrence Prevention Measures: These are managed by dividing them into Short-term, Mid-term, Long-term according to the urgency of the task.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.