Tags: Cloud-Outage

  • Alipay, Taobao, Xianyu Went Dark. Smells Like a Message Queue Meltdown.

    December 05, 2025 in Cloud

    Alipay, Taobao, Xianyu Went Dark. Smells Like a Message Queue Meltdown.

    Dec 4, 2025, Taobao, Alipay, and Xianyu all cratered. Users got charged while orders still showed “unpaid,” a carbon copy of the 2024 Double-11 fiasco.

    Dec 4, 2025, Taobao, Alipay, and Xianyu all cratered. Users got charged while orders still showed “unpaid,” a carbon copy of the 2024 Double-11 fiasco.

  • Cloudflare’s Nov 18 Outage, Translated and Dissected

    November 19, 2025 in Cloud

    Cloudflare’s Nov 18 Outage, Translated and Dissected

    A ClickHouse permission tweak doubled a feature file, tripped a Rust hard limit, and froze Cloudflare’s core traffic for six hours—their worst outage since 2019. Here’s the full translation plus commentary.

    A ClickHouse permission tweak doubled a feature file, tripped a Rust hard limit, and froze Cloudflare’s core traffic for six hours—their worst outage since 2019. Here’s the full translation plus commentary.

  • AWS’s Official DynamoDB Outage Postmortem

    October 24, 2025 in Cloud

    AWS’s Official DynamoDB Outage Postmortem

    AWS finally published the Oct 20 us-east-1 postmortem. I translated the key parts and added commentary on how one DNS bug toppled half the internet.

    AWS finally published the Oct 20 us-east-1 postmortem. I translated the key parts and added commentary on how one DNS bug toppled half the internet.

  • How One AWS DNS Failure Cascaded Across Half the Internet

    October 21, 2025 in Cloud

    How One AWS DNS Failure Cascaded Across Half the Internet

    us-east-1’s DNS control plane faceplanted for 15 hours and dragged 142 AWS services—and a good chunk of the public internet—down with it. Here’s the forensic tour.

    us-east-1’s DNS control plane faceplanted for 15 hours and dragged 142 AWS services—and a good chunk of the public internet—down with it. Here’s the forensic tour.

  • OpenAI Global Outage Postmortem: K8S Circular Dependencies

    December 14, 2024 in Cloud

    OpenAI Global Outage Postmortem: K8S Circular Dependencies

    Even trillion-dollar unicorns can be a house of cards when operating outside their core expertise.

    Even trillion-dollar unicorns can be a house of cards when operating outside their core expertise.

  • Alibaba-Cloud: High Availability Disaster Recovery Myth Shattered

    September 17, 2024 in Cloud

    Alibaba-Cloud: High Availability Disaster Recovery Myth Shattered

    Seven days after Singapore Zone C failure, availability not even reaching 8, let alone multiple 9s. But compared to data loss, availability is just a minor issue

    Seven days after Singapore Zone C failure, availability not even reaching 8, let alone multiple 9s. But compared to data loss, availability is just a minor issue

  • What Can We Learn from NetEase Cloud Music's Outage?

    August 18, 2024 in Cloud

    What Can We Learn from NetEase Cloud Music's Outage?

    NetEase Cloud Music experienced a two-and-a-half-hour outage this afternoon. Based on circulating online clues, we can deduce that the real cause behind this incident was...

    NetEase Cloud Music experienced a two-and-a-half-hour outage this afternoon. Based on circulating online clues, we can deduce that the real cause behind this incident was...

  • Database Deletion Supreme - Google Cloud Nuked a Major Fund's Entire Cloud Account

    May 11, 2024 in Cloud

    Database Deletion Supreme - Google Cloud Nuked a Major Fund's Entire Cloud Account

    Due to an "unprecedented configuration error," Google Cloud mistakenly deleted trillion-RMB fund giant **UniSuper**'s entire cloud account, cloud environment and all off-site backups, setting a new record in cloud computing history!

    Due to an "unprecedented configuration error," Google Cloud mistakenly deleted trillion-RMB fund giant **UniSuper**'s entire cloud account, cloud environment and all off-site backups, setting a new record in cloud computing history!

  • What Can We Learn from Tencent Cloud's Major Outage?

    April 14, 2024 in Cloud

    What Can We Learn from Tencent Cloud's Major Outage?

    Tencent Cloud's epic global outage after Double 11 set industry records. How should we evaluate and view this failure, and what lessons can we learn from it?

    Tencent Cloud's epic global outage after Double 11 set industry records. How should we evaluate and view this failure, and what lessons can we learn from it?

  • From Cost-Reduction Jokes to Real Cost Reduction and Efficiency

    November 29, 2023 in Cloud

    From Cost-Reduction Jokes to Real Cost Reduction and Efficiency

    Alibaba-Cloud and Didi had major outages one after another. This article discusses how to move from cost-reduction jokes to real cost reduction and efficiency — what costs should we really reduce, what efficiency should we improve?

    Alibaba-Cloud and Didi had major outages one after another. This article discusses how to move from cost-reduction jokes to real cost reduction and efficiency — what costs should we really reduce, what efficiency should we improve?

  • What Can We Learn from Alibaba-Cloud's Global Outage?

    November 13, 2023 in Cloud

    What Can We Learn from Alibaba-Cloud's Global Outage?

    Alibaba-Cloud's epic global outage after Double 11 set an industry record. How should we evaluate this incident, and what lessons can we learn from it?

    Alibaba-Cloud's epic global outage after Double 11 set an industry record. How should we evaluate this incident, and what lessons can we learn from it?