Skip to content

GitHub's latest outage blamed on policy error

The "cascade" of failures occurred just as Cursor launched its alternative code hosting platform.

GitHub's latest outage blamed on policy error
Image credit: https://unsplash.com/@jonmsailer

A series of cascading failures sparked by network saturation at Microsoft's Central US data centre were responsible for an almost eight-hour GitHub outage this week.

An incident report posted after the event said web/API error rates had reached 20% during its peak, with the issue persisting after capacity problems were resolved due to a “retry storm” at another data centre.

It said: “Archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com.”

Several scraping attacks on GitHub codeload endpoints during the incident also complicated recovery.

What happened?

This content is for members only

Subscribe
Add The Stack on Google