GitHub Releases August Availability Report; Azure Migration and Architectural Improvements Underway
Key point
GitHub released its August availability report, detailing the status of its Azure migration and the causes of outages along with improvement plans for key services such as Actions and Copilot.
Details
GitHub released its August 2026 Availability Report, providing details on the status of its infrastructure migration to Azure and architectural improvements. Currently, the Azure conversion rate for read traffic is at 60.4% for migration services, 64.3% for the monolith, and 54% for Git reads. By separating the authentication core, approximately 1 million queries per second were reduced, and through query hygiene improvements, 120,000 queries per second and 59,000 seconds per hour of waste were eliminated.
Major Incidents and Root Cause Analysis
In August, multiple incidents occurred in core services such as GitHub Actions, Issues/PR, and Copilot.
- GitHub Actions (8/6, approx. 10 hours 42 minutes): Affected at least 74 organizations. During routine deployment, pod replacement temporarily reduced capacity at one site, and traffic shifting caused the remaining sites to exceed their limits. CPU throttling and OOM restarts in the Service mesh sidecar caused cascading failures, and a latent bug in the job-assignment path caused runners to enter infinite retries, delaying recovery.
- Issues/PR and Widespread Services (8/17, approx. 6 hours 44 minutes): Affected approximately 29K organizations and a total of 4.8M requests. A sudden surge in traffic exceeded the data center load balancer limits, and the Service-mesh sidecar failed to scale up after reaching its concurrency limit. A client retry bug caused a spike in traffic to internal authentication endpoints, creating a cascading effect that delayed the recovery of the Copilot Token Service.
- Copilot Cloud Agent (8/20, approx. 10 hours 40 minutes): Affected at least 54 organizations. Due to a DB region provider outage, status/results updates were delayed by up to 60–90 minutes. The tasks themselves executed normally, but DB delays caused a backlog in the streaming processor.
- GitHub Actions (8/26, approx. 2 hours 53 minutes): Partially affected 386 organizations. Due to insufficient shared infrastructure relative to Actions growth, event bursts occurred, saturating the DB primary. The absence of an automatic circuit breaker necessitated manual throttling.
- Copilot Kimi K3 Model (8/27, approx. 2 hours 50 minutes): Serving degradation from the upstream model provider caused 63.3% of Kimi K3 requests to fail. This affected only users of that model, while other models and Auto settings functioned normally.
Improvement and Response Plans
To prevent such incidents, GitHub is pursuing measures such as securing headroom for Service mesh ingress and the actions service, and enabling autoscaling. It is strengthening logic to prevent capacity reduction during deployments and improving saturation and database-proxy monitoring. Additionally, it plans to enhance system stability by fixing client retry behavior, strengthening LB capacity monitoring, and reinforcing regional failover. For Copilot, it is removing storage configs that cause failover delays and transitioning to a streaming structure resilient to latency.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.