Incident Report for GitHub outage on 2026-08-17
摘要
GitHub 于 2026 年 8 月 17 日 13:28–21:15 UTC 发生大规模故障,持续 7 小时 47 分,影响 Issues、Pull Requests、API、Actions、Copilot 等服务。峰值时 Web/API 错误率约 20%,归档与原始内容下载错误率约 50%,SAML/OIDC 认证、SCIM、Team Sync 及依赖 GitHub.com 公共工作流定义的 GHEC 数据驻留 Actions 也受影响。根因是中央美国区负载均衡器网络饱和,由 Istio sidecar 达到并发限制且自动扩缩容策略配置错误引发,级联导致 4 个 HAProxy 节点耗尽流限制,认证路径降级;VS Code 对单个内部端点延迟响应的重试 bug 将流量放大约 10 倍,拖慢 Copilot Token Service 恢复。部分流量被转移至北弗吉尼亚成功服务。恢复措施包括暂停 HAProxy、通过 PR 降低网关重试逻辑、在负载均衡器以 403 阻断 Copilot Token 请求并逐步恢复流量。codeload 端点遭爬虫攻击也妨碍了恢复。后续预防措施包括修正自动扩缩容策略、审计 Istio 限制、审查重试与退避行为、修复 VS Code 重试问题、改进负载均衡器容量监控与区域故障转移保障。
荐读理由
从这份复盘可直接抄走几个运维教训:自动扩缩容策略必须把服务网格 sidecar 的并发限制算进去,否则流量峰值会打穿负载均衡器;客户端乐观重试会放大 10 倍流量拖慢恢复,重试限流和退避值得你提前设计。
原文

Get email notifications whenever GitHub creates, updates or resolves an incident.
Get text message notifications whenever GitHub creates or resolves an incident.
Get incident updates and maintenance status messages in Slack.
By subscribing you acknowledge our Privacy Policy. In addition, you agree to the Atlassian Cloud Terms of Service and acknowledge Atlassian's Privacy Policy.
Get webhook notifications whenever GitHub creates an incident, updates an incident, resolves an incident or changes a component status.
Visit our support site.
Get the Atom Feed or RSS Feed.
Incident with GitHub.com
Incident Report for GitHub
Resolved
On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.
Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.
The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.
The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed.
Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.
Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.
To prevent recurrence, our follow-up actions include:
Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.
Auditing Istio request, concurrency, and scaling limits across affected services.
Reviewing retry limits and backoff behavior across gateways and clients.
Addressing the VS Code retry behavior that amplified Copilot token traffic.
Improving load-balancer capacity monitoring and regional failover safeguards.
Posted Aug 17, 2026 - 21:15 UTC
Update
We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.
Posted Aug 17, 2026 - 20:45 UTC
Update
Issues is operating normally.
Posted Aug 17, 2026 - 20:22 UTC
Update
We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.
Posted Aug 17, 2026 - 20:08 UTC
Update
We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.
Posted Aug 17, 2026 - 19:13 UTC
Update
API Requests is operating normally.
Posted Aug 17, 2026 - 19:01 UTC
Update
API Requests is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 18:48 UTC
Update
The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.
Posted Aug 17, 2026 - 18:23 UTC
Update
We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.
Posted Aug 17, 2026 - 18:11 UTC
Update
Issues is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 17:36 UTC
Update
We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.
Posted Aug 17, 2026 - 17:34 UTC
Update
Git Operations is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 17:30 UTC
Update
The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.
Posted Aug 17, 2026 - 16:59 UTC
Update
We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.
Posted Aug 17, 2026 - 16:36 UTC
Update
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.
Posted Aug 17, 2026 - 16:16 UTC
Update
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.
Posted Aug 17, 2026 - 15:42 UTC
Update
Webhooks is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 15:40 UTC
Update
Git Operations is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 15:21 UTC
Update
Pages is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 15:10 UTC
Update
API Requests is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 15:01 UTC
Update
Webhooks is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:58 UTC
Update
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.
Posted Aug 17, 2026 - 14:58 UTC
Update
Actions is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:58 UTC
Update
Pull Requests is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:54 UTC
Update
Issues is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:49 UTC
Update
Pull Requests is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:45 UTC
Update
Copilot is experiencing degraded availability. We are continuing to investigate.
Posted Aug 17, 2026 - 14:31 UTC
Update
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.
Posted Aug 17, 2026 - 14:24 UTC
Update
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.
Posted Aug 17, 2026 - 14:04 UTC
Update
Pull Requests is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 13:58 UTC
Update
Issues is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 13:46 UTC
Update
We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available
Posted Aug 17, 2026 - 13:45 UTC
Update
Webhooks is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 13:44 UTC
Update
Actions is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 13:42 UTC
Update
API Requests is experiencing degraded performance. We are continuing to investigate.
Posted Aug 17, 2026 - 13:41 UTC
Investigating
We are investigating reports of impacted performance for some GitHub services.
Posted Aug 17, 2026 - 13:40 UTC
This incident affected: Git Operations, Webhooks, API Requests, Issues, Pull Requests, Actions, Pages, and Copilot.
这条对你有帮助吗?