Devsisters' Incident Response Principles and Methods
Key point
Incidents are handled with rapid recovery as the top priority, supported by a tiering, logging, and retrospective system.
Details
Devsisters sets service normalization as the top priority in incident response, emphasizing prompt recovery over root-cause identification. Anyone capable of helping should pitch in immediately regardless of their job role, and if judgment is difficult, they should ask for help without delay.
To prepare for incident response, everyday readiness is needed, such as carrying a laptop, having development/operations environments set up, and having tethering options available. Alarms must be configured so they are always received, and every alarm should be acknowledged even if only briefly; unnecessary alarms should have their conditions adjusted or be reassigned to a different tier.
Alarms are divided into Tier 0, 1, 2 based on FRT (First Response Time). Tier 0 covers issues that customers directly experience or that pose a direct risk to service stability, requiring the fastest response. Tier 1 covers technical issues that don't immediately affect customers but could escalate into Tier 0, while Tier 2 covers issues with low direct impact but that require intervention within a set timeframe, such as certificate expiration, policy violations, or cost alarms.
When an actual incident occurs, the team should be notified immediately, and a response team of at least 2 people should be formed. Roles are split into Commander and Scribe: the Commander sets direction and drives decision-making, while the Scribe thoroughly records the cause, hypotheses, actions taken, and a chronological timeline.
The incident response channel is opened via Datadog Incident, and the main communication is kept in Slack messages so it can be tracked later. When an incident ends, the end time, estimated cause, actions taken, and postmortem schedule should be clearly shared, and the status should be updated from Active to Stable or Resolved.
Afterward, a postmortem is conducted to organize the cause, actions, lessons learned, and action items so the same problem doesn't recur. The principle is a blameless retrospective, emphasizing that failures should be treated as opportunities to strengthen the system rather than to blame individuals.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.