Facebook, WhatsApp, Instagram, Oculus Outage
Discover how a routine configuration change led to the 2021 Facebook global outage through cascading failures and DNS withdrawal. Learn essential System Design lessons on avoiding automation pitfalls, ensuring operational readiness, and implementing simple, robust contingency plans for core services.
We'll cover the following...
In October 2021, Facebook experienced a six-hour global outage affecting related services, including Messenger, WhatsApp, Instagram, and Oculus. The New York Times described the event with the headline: “Gone in Minutes, Out for Hours: Outage Shakes Facebook.” Estimates suggest the outage cost Facebook about $100 million in revenue and billions in market value. The following sequence of events led to the outage:
The sequence of events
A chain reaction of technical failures led to the total blackout:
Routine maintenance: An automated system attempted to assess spare capacity on Facebook’s backbone network.
Configuration error: A faulty command accidentally disconnected all data centers from the backbone network. An automated audit tool, designed to catch such errors, failed to detect the bug.
DNS health checks: Facebook’s authoritative DNS servers monitor network health. Because they could not reach the internal data centers, they executed a fail-safe protocol to stop advertising their presence on the internet by withdrawing their
routes.BGP Border Gateway Protocol Resolution failure: With routes withdrawn, the authoritative DNS servers became unreachable. Public DNS resolvers (like Google or Cloudflare) could not refresh their records. As cached entries for www.facebook.com timed out, the domain became unresolvable globally.
Total outage: Consequently, no traffic could reach Facebook or its subsidiaries.
The slides below visually depict these events: