AI Briefing
KO

Cookie Run: Kingdom AWS AZ Outage Analysis

·2023.02.16 09:00

Key point

A cooling system failure in the AWS Tokyo region caused 6 CockroachDB nodes to fail simultaneously, resulting in partial data loss, which was analyzed.

1 / 2

Details

In February 2021, a cooling system failure at an AWS Tokyo region data center caused servers to overheat, rendering multiple nodes inoperable. As a result, 6 of Devsisters' CockroachDB nodes sequentially failed over 22 minutes.

CockroachDB is a distributed database based on the Raft algorithm that maintains service as long as a majority of nodes remain operational. At the time, Devsisters had set the Replication Factor to 7, designed to withstand up to 3 node failures, but 6 nodes failed simultaneously, far exceeding expectations.

As a result, out of a total of 25,000 Ranges, 2 Data Ranges and 32 System Ranges were lost. This occurred because Locality settings accounting for Multi-AZ had not been applied, as the team trusted the cloud service's reliability and had only prepared for Instance Retirement scenarios.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.