Atmos Pro Logo

Atmos Pro

ProductPricingDocsBlogChangelog
Create Workspace

Incident History

Incident reports, postmortems, and status updates from Atmos Pro and its infrastructure dependencies.

← Back to System Status·RSS Feed
Our IncidentsDependenciesSecurity
Atmos Pro Logo

Atmos Pro

The fastest way to deploy your apps on AWS with Terraform and GitHub Actions.

GitHubTwitterLinkedInYouTubeSlack

For Developers

  • Quick Start
  • Example Workflows
  • Atmos Documentation
  • Register for Office Hours
  • Join the Slack Community

Enterprise

  • Trust Center
  • Procurement
  • Pricing
  • Contact Sales

Company

  • About Cloud Posse
  • Security
  • Blog
  • Media Kit
  • Try our Newsletter

Legal

  • SaaS Agreement
  • Terms of Use
  • Privacy Policy
  • Disclaimer
  • Cookie Policy

© 2026 Cloud Posse, LLC. All rights reserved.

Checking status...

Incident updates from the upstream infrastructure Atmos Pro depends on — pulled directly from the status pages of services like GitHub, Vercel, Inngest, and Resend.

GitHubresolvedSep 4, 2026 at 22:02 UTC

Degradation in repos contents API

Sep 4, 22:23 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Sep 4, 22:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedSep 4, 2026 at 20:39 UTC

Disruption with Copilot Code Review

Sep 4, 22:26 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Sep 4, 22:25 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Sep 4, 21:54 UTC
Update - We are applying the mitigation and expect recovery within approximately 30 minutes.

Sep 4, 20:57 UTC
Update - Some users may be experiencing failures when using Copilot code review. We have identified the root cause and are working on a mitigation.

Sep 4, 20:39 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendSep 4, 2026 at 19:34 UTC

Delays in background processes are affecting webhooks, domain verifications, and other important features

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • General API (Operational)
  • Webhooks (Operational)
  • Dashboard (Operational)
View →
VercelresolvedSep 4, 2026 at 19:11 UTC

Elevated latency in VCR, Blob, and Sandbox API

Sep 4, 20:27 UTC
Resolved - Blob has fully recovered. This incident has been resolved.

Sep 4, 20:19 UTC
Update - Elevated latency has recovered for VCR and Sandbox. We are continuing to investigate elevated latency affecting Blob and will provide another update shortly.

Sep 4, 20:05 UTC
Monitoring - The fix has been rolled out and we are seeing signs of recovery.

Sep 4, 19:55 UTC
Update - The fix is currently being rolled out. We'll provide another update once the rollout is complete and we're beginning to see signs of recovery.

Sep 4, 19:33 UTC
Update - We are continuing to work on a fix for this issue.

Sep 4, 19:33 UTC
Identified - The issue has been identified and a fix is being implemented.

Sep 4, 19:11 UTC
Investigating - We've identified an issue where some customers may experience increased latency for API requests to VCR, Blob, and Sandbox. We are currently investigating this issue. We will provide additional updates as they become available.

View →
InngestSep 4, 2026 at 19:10 UTC

Degraded Function Execution

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
GitHubresolvedSep 3, 2026 at 14:17 UTC

Incident with Grok Copilot AI Model Provider

Sep 3, 17:11 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Sep 3, 17:11 UTC
Update - The issues with our upstream model provider have been resolved, and Grok models are once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability, but mitigation is complete.

Sep 3, 14:22 UTC
Update - The Grok 4.5 model has degraded availability as well. We are working with the upstream provider to resolve the issue.

Sep 3, 14:20 UTC
Update - We are experiencing degraded availability for the Grok 4.6 model in Copilot Chat, VS Code and other Copilot products. This is due to an issue with an upstream model provider. We are working with them to resolve the issue.

Sep 3, 14:17 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
ResendSep 3, 2026 at 10:45 UTC

Delay in contact webhooks delivery

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Webhooks (Operational)
View →
VercelresolvedSep 1, 2026 at 19:59 UTC

Increased deployment failures

Sep 1, 20:58 UTC
Resolved - This incident has been resolved. Builds have recovered.

Sep 1, 20:49 UTC
Update - We are continuing to monitor for any further issues.

Sep 1, 20:45 UTC
Monitoring - We are seeing recovery in Builds and are continuing to monitor.

Sep 1, 20:37 UTC
Identified - We have identified the issue and are working on a fix. We are seeing recovery as the fix is deployed. We will provide additional updates as they become available.

Sep 1, 20:27 UTC
Update - We are actively working to fix an issue causing some deployments that use IAD1 Function regions or Routing Middleware to fail. We will provide additional updates as they become available.

Sep 1, 19:59 UTC
Investigating - We've identified an issue where some deployments that use IAD1 Function regions or Routing Middleware are failing. We are investigating and will provide more information as it becomes available.

View →
GitHubresolvedSep 1, 2026 at 15:00 UTC

Delays in commit processing

Sep 1, 16:01 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Sep 1, 16:00 UTC
Update - Time to update pull request diffs have improved to normal thresholds.

Sep 1, 15:00 UTC
Update - Diffs in the PR view may be stale for several minutes. We are investigating and scaling up resources.

Sep 1, 15:00 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
GitHubresolvedAug 31, 2026 at 09:15 UTC

Elevated rate of errors for OpenAI models provided by Copilot

Aug 31, 09:58 UTC
Resolved - Between 08:37 and 09:41 UTC on August 31, 2026, GitHub Copilot experienced degradation affecting several GPT models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, and the gpt-5.6 family (Luna, Sol, and Terra). Users encountered elevated error rates and interrupted streaming responses. Other models were not affected.

The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.

Aug 31, 09:58 UTC
Update - The issues with our upstream model provider have been resolved, and gpt-5.3-codex, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5, and the gpt-5.6 family of models are once again available in Copilot products and IDE surfaces.

We will continue monitoring to ensure stability, but mitigation is complete.

Aug 31, 09:51 UTC
Monitoring - The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

Aug 31, 09:48 UTC
Update - One of our model providers has confirmed an incident on their end. We have provided them details to help identify the issue. We are starting to see recovery.

Aug 31, 09:22 UTC
Update - Copilot is experiencing a higher rate of errors for OpenAI models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4, and the gpt-5.6 family of models. Other models are not impacted.

Aug 31, 09:15 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
GitHubresolvedAug 27, 2026 at 10:04 UTC

Incident with Copilot AI Model Providers

Aug 27, 12:12 UTC
Resolved - On August 27th, 2026, between approximately 09:20 and 12:14 UTC, the Copilot service experienced a degradation of the Kimi K3 model due to an issue with our upstream provider. Users encountered elevated error rates when using Kimi K3. No other models were impacted.

The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

Aug 27, 12:12 UTC
Update - The issues with our upstream model provider have been mitigated, and Kimi K3 is once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability.

Aug 27, 11:58 UTC
Update - Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.

Aug 27, 10:43 UTC
Update - We are experiencing degraded availability for the Kimi K3 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Aug 27, 10:04 UTC
Investigating - We are investigating reports of degraded availability for Copilot AI Model Providers

View →
GitHubresolvedAug 26, 2026 at 23:37 UTC

Disruption with GitHub Billing

Aug 27, 19:44 UTC
Resolved - On August 26, 2026, between 20:40 UTC and 00:51 UTC on August 27, GitHub Billing experienced degraded performance affecting billing budget pages and GitHub Copilot CLI sessions. Affected customers encountered failed budget page loads or failures when starting or continuing CLI sessions. We confirmed this impact for a small number of customers (<1%).

This was caused by a concentrated workload that created processing delays in our data storage layer. Automated retries increased the load and prolonged the degradation. We mitigated the incident by rebalancing traffic within our infrastructure.

We are improving workload isolation, retry behavior, and detection of concentrated load to reduce the likelihood of recurrence and shorten our time to detect and mitigate similar incidents.

Aug 27, 17:58 UTC
Update - No material change since the previous update. Service conditions remain stable following the mitigation, and we have not observed any further customer impact. We are actively monitoring the service while implementing targeted fixes to address the underlying root cause.

Aug 27, 16:20 UTC
Update - Our mitigation continues to hold, and service conditions remain stable. We are continuing to investigate the concentrated workload responsible for the issue and are preparing additional preventative improvements. We have not identified a material change in customer impact since the previous update. We will provide another update as the investigation progresses.

Aug 27, 14:49 UTC
Update - Our mitigation is still holding as we continue to investigate to find the root cause.

Aug 27, 01:35 UTC
Update - We are continuing to monitor the mitigation that we have applied for the billing page disruption.

Aug 27, 00:31 UTC
Update - We've applied a mitigation to unblock Copilot usage and have observed recovery for this particular impact. We're continuing to investigate and apply mitigations for the billing page disruption while monitoring to ensure Copilot remains recovered.

Aug 26, 23:42 UTC
Update - We are currently investigating increased errors with billing services. Customers may observe failed billing budget page loads, and users of the Copilot CLI may observe failures starting or continuing sessions.

Aug 26, 23:37 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 26, 2026 at 22:57 UTC

Incident with Actions and Pull Requests

Aug 27, 00:26 UTC
Resolved - On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed.

Background jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information.

We reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations.

We are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load.

Aug 27, 00:26 UTC
Monitoring - The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability.

Aug 27, 00:25 UTC
Update - We confirmed full recovery beginning at 23:58 UTC. Actions workflow runs and pull request merges are operating normally. We will now resolve the incident while continuing to monitor service health.

Aug 27, 00:01 UTC
Update - We've applied mitigations and are seeing recovery in Actions workflow runs and blocked pull request merges. We're continuing to monitor for sustained health of merge commit creates before resolving.

Aug 26, 22:57 UTC
Update - We are investigating elevated delays and timeouts affecting Actions workflow runs triggered by pull request events. 20% of actions runs have delayed starts of more than 5 minutes and up to 4% of runs failed to trigger. We are actively working on mitigation and will provide updates as we learn more.

Aug 26, 22:56 UTC
Investigating - We are investigating reports of degraded performance for Actions and Pull Requests

View →
GitHubresolvedAug 26, 2026 at 15:12 UTC

Incident with Actions

Aug 26, 18:01 UTC
Resolved - On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system.

At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC.

3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs.

Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27.

Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs.

Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.

Aug 26, 18:00 UTC
Update - All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.

Aug 26, 17:54 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Aug 26, 17:32 UTC
Update - We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.

Aug 26, 16:50 UTC
Update - We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.

Aug 26, 16:49 UTC
Update - Pages is operating normally.

Aug 26, 16:14 UTC
Update - We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

Aug 26, 15:48 UTC
Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

Aug 26, 15:23 UTC
Update - We've identified an issue with a database primary and are failing over to a replica immediately

Aug 26, 15:12 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Aug 26, 15:11 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
GitHubresolvedAug 26, 2026 at 15:09 UTC

Disruption with some GitHub services

Aug 26, 16:07 UTC
Resolved - Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/y1t7p9fzrlj2

Aug 26, 15:09 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedAug 26, 2026 at 03:06 UTC

Failures logging in with Vercel CLI

Aug 26, 04:28 UTC
Resolved - This incident has been resolved.

Aug 26, 04:28 UTC
Identified - The team has identified the source of the issue and is testing a fix.

Aug 26, 03:15 UTC
Update - We are investigating errors logging in with Vercel CLI (vc login).

Aug 26, 03:06 UTC
Investigating - We are investigating errors logging in with Vercel CLI (vc login).

View →
InngestAug 26, 2026 at 00:31 UTC

Connect Message Routing Delays - for Connect users only

Status: Resolved

Between August 26 00:31 to 01:13 UTC, a subset of customers using Inngest Connect experienced delayed and failed function runs due to an internal networking issue that disrupted connectivity between Connect services. We reverted the change and restored service. Although Connect buffers worker responses, stalled internal API calls delayed acknowledgements and subsequent replies. In some cases, request forwarding failed even after a buffered response had been received. We fixed the root cause, improved Connect’s failure handling, and enhanced our internal observability so we can detect and respond more quickly to similar problems in the future. Affected customers should review failed Connect runs during the incident window and replay them after confirming that doing so is safe and idempotent.
View →
VercelresolvedAug 25, 2026 at 19:27 UTC

Elevated build initialization times

Aug 25, 20:05 UTC
Resolved - This incident has been resolved.

Aug 25, 19:57 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Aug 25, 19:37 UTC
Identified - The issue has been identified and a fix is being implemented.

Aug 25, 19:27 UTC
Investigating - We're investigating elevated initialization times for some builds. We will provide updates as we learn more.

View →
InngestAug 24, 2026 at 20:38 UTC

Connect Message Routing Delays - for Connect users only

Status: Resolved

Between August 24 at 23:38 UTC and August 25 at 00:08 UTC, a subset of customers using Inngest Connect experienced delayed and failed function runs due to an internal networking issue that disrupted connectivity between Connect services. We reverted the change and restored service. Although Connect buffers worker responses, stalled internal API calls delayed acknowledgements and subsequent replies. In some cases, request forwarding failed even after a buffered response had been received. We fixed the root cause, improved Connect’s failure handling, and enhanced our internal observability so we can detect and respond more quickly to similar problems in the future. Affected customers should review failed Connect runs during the incident window and replay them after confirming that doing so is safe and idempotent.
View →
GitHubresolvedAug 24, 2026 at 13:56 UTC

Actions delays in starting runs

Aug 24, 14:34 UTC
Resolved - On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright.

The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC.

To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.

Aug 24, 14:26 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Aug 24, 14:22 UTC
Update - Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.

Aug 24, 13:56 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
GitHubresolvedAug 24, 2026 at 07:12 UTC

Elevated errors on Fable 5 due to upstream provider

Aug 24, 07:58 UTC
Resolved - On August 24th, 2026, between approximately 06:35 and 07:25 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates when using Claude Fable 5, with requests sometimes failing mid-response. No other models were impacted.

The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

Aug 24, 07:12 UTC
Update - We are experiencing degraded availability for the Fable model in Copilot products and IDE surfaces. This is due to an issue with the upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Aug 24, 07:12 UTC
Investigating - We are investigating reports of degraded availability for Copilot AI Model Providers

View →
GitHubresolvedAug 21, 2026 at 14:00 UTC

Degraded Git Operations over SSH

Aug 21, 14:00 UTC
Resolved - On August 21, 2026, between 14:00 and 14:07 UTC, dotcom Git operations over SSH were degraded. Successful Git operations over SSH fell by more than 95% for during the peak impact window, making clone, fetch, or push over SSH effectively unavailable to most users for approximately four minutes. Git operations over HTTPS were not affected.

The incident was caused by a software defect in our load-balancing infrastructure that was triggered by a configuration change. The defect only occurred when connections passed through multiple layers of load balancers running the new configuration, which meant it was not detected during canary testing.

We mitigated the incident by rolling back the configuration change.

We are adding regression coverage for multi-layer load-balancer configurations and improving monitoring and alerting for Git operations over SSH to reduce our time to detection and mitigation of similar issues in the future.

View →
InngestAug 20, 2026 at 22:35 UTC

We are currently experiencing a delay with function runs being visible in the dashboard and have scaled up processing to work through the backlog.

Status: Resolved

Run history delay has caught up and the incident is now resolved.

Affected components
  • Observability (Operational)
View →
VercelresolvedAug 20, 2026 at 17:46 UTC

New container functions are failing

Aug 20, 18:34 UTC
Resolved - This incident has been resolved.

Aug 20, 18:23 UTC
Monitoring - We have deployed a fix. Newly deployed container functions are working as expected.

Aug 20, 17:46 UTC
Identified - We've identified an issue where newly deployed container functions may return 5XX errors.

Customers can use a previously working container function deployment as a temporary mitigation. We are working on a fix and will provide additional updates as they become available.

View →
InngestAug 20, 2026 at 15:04 UTC

Delays in Function Execution

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
GitHubresolvedAug 20, 2026 at 14:43 UTC

Intermittent failures creating agent tasks

Aug 21, 00:37 UTC
Resolved - Between 13:57 UTC on August 20 and 00:37 UTC on August 21, 2026, some users of the Copilot Cloud Agent experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. The agent tasks themselves continued to run and complete during this time; only the visibility of their status was delayed.

The cause was a regional outage in a third-party cloud database service that Copilot uses to store agent task status. We failed over the affected database to a healthy region, added processing capacity to work through the backlog, and restored normal operation once the underlying service recovered. No task data was lost during the incident.

To prevent repetition of similar incidents, we are removing the database configuration that made us vulnerable to this regional outage and improving our database failover procedures.

Aug 20, 20:37 UTC
Update - We are seeing gradual recovery in Copilot Cloud Agent task status visibility as we deploy a fix for the root cause. Session output remains delayed by approximately one hour while remediation continues.

Aug 20, 19:35 UTC
Update - We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility. Session output continues to be delayed by approximately 1 hour as our remediation steps take effect.

Aug 20, 18:45 UTC
Update - We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and expect this to take effect within the next hour.

Aug 20, 18:04 UTC
Update - We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.

Aug 20, 17:32 UTC
Update - We are observing gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.

Aug 20, 17:05 UTC
Update - We are seeing signs of recovery for Copilot Cloud Agent task status visibility, but this recovery is slower than anticipated. We are pursuing additional mitigating measures to accelerate recovery.

Aug 20, 16:14 UTC
Update - Users are experiencing delays when starting tasks using Copilot Cloud Agent and are not be able to see the status of these tasks. Copilot Cloud Agent tasks are still being completed. We have identified the cause of the issue and are putting mitigations in place to return service to normal levels. We will provide another update about the expected recovery time shortly.

Aug 20, 15:41 UTC
Update - We are experiencing issues with Copilot Cloud Agent tasks, resulting in newly started tasks not properly displaying on-going progress. These Copilot Cloud Agent tasks are still being completed correctly but lack proper visibility. We are actively investigating the issue and will provide updates as we learn more.

Aug 20, 15:01 UTC
Update - We have identified the problematic component and are working to fail over to a healthy instance. Further updates will be provided as we perform mitigations.

Aug 20, 14:51 UTC
Update - Users may experience delays when starting tasks using Copilot Cloud Agent. We are actively investigating the issue and will provide updates as we learn more.

Aug 20, 14:43 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedAug 20, 2026 at 11:14 UTC

Logs unavailable to query in Dashboard

Aug 20, 11:25 UTC
Resolved - The system is operating normally and fully recovered.

Aug 20, 11:21 UTC
Monitoring - We identified the source of the failure and implemented a fix. Runtime logs can be queried from Dashboard again. We are continuing to monitor.

Aug 20, 11:14 UTC
Investigating - We are investigating a failure to query runtime logs in the Dashboard. Log ingestion and querying build logs are unaffected.

View →
ResendAug 20, 2026 at 05:22 UTC

Increased latency across API endpoints

Status: Resolved

We have resolved the underlying issue and service latency has been recovered.

Affected components
  • General API (Operational)
View →
ResendAug 19, 2026 at 10:01 UTC

Contact webhook events are delayed in delivery

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Webhooks (Operational)
View →
InngestAug 19, 2026 at 04:41 UTC

Delayed function executions for a subset of customers

Status: Resolved

Execution latencies are back to normal across all customers, and the incident is now resolved.

Affected components
  • Function execution (Operational)
View →
VercelresolvedAug 18, 2026 at 23:00 UTC

Elevated errors for Workflow streams

Aug 18, 23:00 UTC
Resolved - From 10:57pm to 11:43pm UTC, a percentage of Workflows streams calls failed. Workflow steps that had a failing stream call would retry, and if the step hit the max retries, the workflow failed. Workflow streams have recovered and are working correctly now.

View →
GitHubresolvedAug 18, 2026 at 09:36 UTC

Incident with Actions

Aug 18, 10:23 UTC
Resolved - On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to run jobs on Actions Larger Runners and were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API.

These issues were caused by failures in backend requests resolving essential metadata for starting Larger Runner workflow runs and for reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents that had been triggered by this operation.

We mitigated the issues by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. The relevant service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.

Aug 18, 09:36 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 18, 2026 at 07:40 UTC

Intermittent failures in runner group and runner-related permissions pages

Aug 18, 11:42 UTC
Resolved - On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API.

The issue was caused by failures in backend requests reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents triggered by this operation.

The impact was mitigated by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. This service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.

Aug 18, 11:24 UTC
Update - We have applied a mitigation and are seeing recovery signals. We will continue monitoring recovery and providing updates.

Aug 18, 10:41 UTC
Update - We have identified the source of a communication issue between Actions services and are working toward mitigation. Customers may experience failure to load runner groups and runner-related permissions issues when using Larger Runners.

Aug 18, 07:40 UTC
Monitoring - We are investigating reports of failure to load runner groups and runner-related permissions for customers using larger runners.

Aug 18, 07:40 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedAug 17, 2026 at 14:30 UTC

Partial disruption of Observability, Analytics, Firewall, Usage, and Login in Dashboard

Aug 17, 21:25 UTC
Resolved - Services have recovered. We are still working to ingest missing observability data during this period.

Aug 17, 20:13 UTC
Update - Services have recovered. We are still working to ingest missing observability data during this period.

Aug 17, 17:51 UTC
Update - Login with TOTP has completely recovered. We are still working to ingest missing observability data during this period. We will share more information as it becomes available.

Aug 17, 16:46 UTC
Update - We are continuing to monitor for any further issues.

Aug 17, 16:44 UTC
Monitoring - We have implemented a fix for the disruption, and services have started to recover. We will share more information as it becomes available.

Aug 17, 16:06 UTC
Update - We are also investigating an issue where some customers may be unable to sign in to Dashboard with time-based one-time passwords (TOTP). We will share updates as they become available.

Aug 17, 15:09 UTC
Update - We are continuing to investigate this issue. We will share updates as they become available.

Aug 17, 14:30 UTC
Investigating - We are investigating a partial disruption to Observability, Web Analytics, Speed Insights, Firewall, and Usage in Dashboard. We will share updates as they become available.

View →
GitHubresolvedAug 17, 2026 at 13:44 UTC

Incident with GitHub.com

Aug 17, 21:15 UTC
Resolved - On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.

Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.

The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.

The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed.

Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.

Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.

To prevent recurrence, our follow-up actions include:

- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.

- Auditing Istio request, concurrency, and scaling limits across affected services.

- Reviewing retry limits and backoff behavior across gateways and clients.

- Addressing the VS Code retry behavior that amplified Copilot token traffic.

- Improving load-balancer capacity monitoring and regional failover safeguards.

Aug 17, 20:45 UTC
Update - We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.

Aug 17, 20:22 UTC
Update - Issues is operating normally.

Aug 17, 20:08 UTC
Update - We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.

Aug 17, 19:13 UTC
Update - We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.

Aug 17, 19:01 UTC
Update - API Requests is operating normally.

Aug 17, 18:48 UTC
Update - API Requests is experiencing degraded availability. We are continuing to investigate.

Aug 17, 18:23 UTC
Update - The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.

Aug 17, 18:11 UTC
Update - We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.

Aug 17, 17:36 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Aug 17, 17:34 UTC
Update - We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.

Aug 17, 17:30 UTC
Update - Git Operations is experiencing degraded performance. We are continuing to investigate.

Aug 17, 16:59 UTC
Update - The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

Aug 17, 16:36 UTC
Update - We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.

Aug 17, 16:16 UTC
Update - We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.

Aug 17, 15:42 UTC
Update - We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.

Aug 17, 15:40 UTC
Update - Webhooks is experiencing degraded performance. We are continuing to investigate.

Aug 17, 15:21 UTC
Update - Git Operations is experiencing degraded performance. We are continuing to investigate.

Aug 17, 15:10 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Aug 17, 15:01 UTC
Update - API Requests is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:58 UTC
Update - Webhooks is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:58 UTC
Update - We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.

Aug 17, 14:58 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:54 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:49 UTC
Update - Issues is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:45 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:31 UTC
Update - Copilot is experiencing degraded availability. We are continuing to investigate.

Aug 17, 14:24 UTC
Update - We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.

Aug 17, 14:04 UTC
Update - We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.

Aug 17, 13:58 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Aug 17, 13:46 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Aug 17, 13:45 UTC
Update - We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available

Aug 17, 13:44 UTC
Update - Webhooks is experiencing degraded performance. We are continuing to investigate.

Aug 17, 13:42 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Aug 17, 13:41 UTC
Update - API Requests is experiencing degraded performance. We are continuing to investigate.

Aug 17, 13:40 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendAug 16, 2026 at 12:57 UTC

API instability

Status: Resolved

API became unstable from 12:54 to 12:57 UTC, returning 5xx for approximately 7% of the requests. We have resolved the underlying issue and service has been resumed.
View →
VercelresolvedAug 14, 2026 at 13:08 UTC

Dynamic API routes returned 404 errors for some deployments

Aug 14, 13:08 UTC
Resolved - This incident has been resolved. Between August 12, 17:12 UTC and August 13, 20:33 UTC, some Vercel deployments could return 404 errors from dynamic API routes.

Deployments built on Vercel during this period were affected only if they met all three conditions: used Next.js 16.2 or earlier, used the Pages Router with internationalization (i18n configured in next.config.js), and used dynamic API routes, such as pages/api/[slug].ts or pages/api/[...path].ts.

Deployments built outside Vercel (with vercel deploy --prebuilt) were not affected, and existing deployments not rebuilt during this period were not affected.

If your project matches the conditions above and you deployed during the impact window, redeploy your project. Affected deployments are not fixed automatically; a new deployment is required to restore the affected routes.

View →
ResendAug 13, 2026 at 17:43 UTC

Errors during the domain creation process

Status: Resolved

Between approximately 17:30 and 17:43 UTC, requests to add a new domain failed in the dashboard and through the public API (POST /domains). Verification of domains added during this window was also affected. The issue was identified and fixed, and domain creation has returned to normal. No domains were partially created: affected requests failed immediately, so retrying the domain creation will succeed. Email sending for existing domains was not affected.

Affected components
  • Dashboard (Operational)
  • General API (Operational)
View →
GitHubresolvedAug 13, 2026 at 16:21 UTC

Disruption with GHEC Team Sync

Aug 13, 18:27 UTC
Resolved - On August 13, 2026, from 15:31:21 UTC to 18:27:55 UTC, GitHub Enterprise Cloud team synchronization was degraded for enterprises using personal accounts. Organization teams experienced delays of up to 3 to 13 hours (median 8 hours) when syncing with IdP groups, resulting in delayed access grants or removals for enterprise users across 2.8% of teams.

A temporary change introduced to address a previous issue due to increased usage of this feature remained active after it was intended to be removed, causing synchronization delays during periods of high volume. We removed the temporary change and provisioned additional resources to handle the increased volume.

Aug 13, 18:27 UTC
Update - We have deployed a mitigation. At this time GHEC Team Sync has recovered for enterprises with personal accounts. Teams syncing to IdP groups have returned to their normal cadence.

Aug 13, 16:21 UTC
Update - GHEC Team Sync is currently degraded for enterprises with personal accounts, causing delays when syncing teams to IdP groups. We have identified the cause of the delays and are working on a mitigation. We will provide an update on our progress at 20:00 UTC.

Aug 13, 16:21 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 13, 2026 at 14:46 UTC

Incident with Webhooks

Aug 13, 15:36 UTC
Resolved - Between 14:24 and 14:53 UTC on 13 August 2026, a routine background job to delete an organization overwhelmed a key shared database, causing multiple GitHub services to briefly return elevated errors and slower responses. Most affected was the webhook management API, with smaller impact to Git operations, pull requests, issues, packages, sign-in, and Copilot. Impact cleared on its own at about 14:53 UTC once the job finished; we resolved the incident at 15:36 UTC.

Affected users may have experienced a brief increase in errors and slower responses, primarily when creating, listing, or updating webhooks, with smaller impacts to pull requests, issues, packages, and Git operations. Failures peaked at about 1% for several minutes around 14:37 UTC.

To prevent future incidents, we've already shipped an update that turns on the safer deletion path for organizations, along with caps on deletion holds on databases. Building on these changes, we're auditing all bulk deletion and cleanup jobs that write to shared databases to prevent similar issues in future.

Aug 13, 15:33 UTC
Update - We have temporarily disabled a background job which caused the impact. At this time the impact is fully mitigated.

Aug 13, 15:33 UTC
Monitoring - The degradation affecting Git Operations, Issues, Packages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

Aug 13, 14:58 UTC
Update - We are currently investigating a brief degradation of service for Git operations (specifically pushes), issues, pull requests, package registry, and webhooks between 14:32 and 14:46 UTC. We have identified the source of the degradation and are investigating mitigation strategies to prevent recurrence.

Aug 13, 14:56 UTC
Update - Packages is experiencing degraded performance. We are continuing to investigate.

Aug 13, 14:46 UTC
Update - Git Operations is experiencing degraded performance. We are continuing to investigate.

Aug 13, 14:46 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Aug 13, 14:46 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Aug 13, 14:45 UTC
Investigating - We are investigating reports of degraded performance for Webhooks

View →
GitHubresolvedAug 13, 2026 at 14:43 UTC

Errors with the Fable 5 Model in Copilot

Aug 13, 15:47 UTC
Resolved - On August 13th, 2026, between approximately 14:06 and 15:47 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates, peaking at 43% and averaging 12%. Users who selected Auto or alternative models were unaffected.

The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

Aug 13, 15:47 UTC
Update - The issues with our upstream model provider have been resolved, and Fable 5 is once again available in Copilot products and IDE surfaces.

We will continue monitoring to ensure stability, but mitigation is complete.

Aug 13, 15:23 UTC
Update - We are seeing modest recovery, but are still experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Aug 13, 14:50 UTC
Update - We are experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Aug 13, 14:43 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
GitHubresolvedAug 12, 2026 at 21:39 UTC

Disruption with Login and Release Asset downloads

Aug 12, 22:56 UTC
Resolved - On August 12 and 13, 2026, some anonymous (logged-out) requests to github.com experienced HTTP 5xx errors when loading pages like the sign-in page, and when downloading release assets, due to an unusual traffic pattern that repeatedly overloaded a part of our infrastructure that serves these types of requests. There were three windows of impact: (1) August 12 from 16:34 to 18:34 UTC, with an average error rate of 16.16% that peaked at 28.6%; (2) August 12 from 19:00 to 22:56 UTC, with an average error rate of 16.55% that peaked at 24.18%; and (3) August 13 from 06:19 to 08:05 UTC, with an average error rate of 2.01% that peaked at 7.49%.
Requests from signed-in users were unaffected.

We mitigated the incidents by applying traffic controls at our network edge that limited any requests matching the pattern identified previously, thereby preventing overload on our systems.

Since these incidents occurred, we have tightened our monitoring systems to alert server-side errors that affect logged-out traffic. We are also working to further strengthen our edge protections and reduce the time to detect and mitigate similar incidents.

Aug 12, 22:22 UTC
Update - We have identified the root cause and are working on mitigation. Errors on the login page and downloading release assets have decreased, but we are not fully mitigated. We will continue to provide updates.

Aug 12, 21:43 UTC
Update - We are investigating issues with Login and when downloading Release Assets. We will continue to keep users updated on progress towards mitigation.

Aug 12, 21:39 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 12, 2026 at 16:16 UTC

Incident with Pull Requests and Issues

Aug 12, 16:41 UTC
Resolved - Between 16:03 and 16:29 UTC on August 12, some users encountered errors when viewing pull requests, issues, and search results. During this period, about 1.9% of Pull Request requests and 0.9% of Issues requests failed. During a database migration, two indexes were removed while application settings still referenced them, causing affected requests to fail. We detected the issue after the migration reached one database shard and before it progressed to the remaining shards. We restored service by disabling both settings. We are improving safeguards around database migrations and application configuration to prevent similar mismatches from causing errors.

Aug 12, 16:38 UTC
Update - We identified the source of errors affecting Pull Requests, Issues, and Search on GitHub.com and have applied a mitigation. A database index hint was referencing an index that had been removed by a recent migration, causing query failures for some users. We disabled the problematic configuration and are seeing recovery across affected services. We are continuing to monitor to confirm full resolution.

Aug 12, 16:35 UTC
Monitoring - The degradation affecting Issues and Pull Requests has been mitigated. We are monitoring to ensure stability.

Aug 12, 16:24 UTC
Update - We are investigating reports of errors affecting Pull Requests and Issues on GitHub.com. Some users may encounter 500 errors when loading pull request and issue pages. Our engineering teams are actively investigating the root cause, which appears to be related to a database infrastructure issue. We will provide an update as soon as we have more information.

Aug 12, 16:16 UTC
Investigating - We are investigating reports of degraded performance for Issues and Pull Requests

View →
GitHubresolvedAug 11, 2026 at 14:50 UTC

Incident with GraphQL API Requests

Aug 11, 20:06 UTC
Resolved - On August 11, 2026, between 14:00 UTC and 16:00 UTC the GraphQL API service was degraded and customers in saw higher than normal timeouts. On average, the timeout rate was 0.06% and peaked at 0.14% of requests routing to the service.

This was due to increased utilization at one of our sites which caused resource contention across our dependencies, leading to an increase in timeouts for GraphQL requests. We mitigated the incident by increasing capacity to alleviate the capacity bottleneck.

We are working to improve our monitoring so that we can proactively reduce the impact of high consumption requests in addition to scaling up; Additionally, we will improve our time to detection and mitigation of issues like this one in the future.

Aug 11, 20:06 UTC
Update - We have identified and mitigated increased error rates affecting GraphQL API requests. A fix to increase service capacity has been deployed and error rates have returned to normal levels. We are resolving this incident.

Aug 11, 16:49 UTC
Monitoring - The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

Aug 11, 16:49 UTC
Update - We have returned to a healthy baseline on GraphQL API requests. We will continue to work on investigations into the errors seen during this incident.

Aug 11, 15:28 UTC
Update - We are investigating reports of a small increase in error rates affecting GraphQL API requests. We are working on increasing capacity and continue to investigate the increased errors. We will provide another update when we have more information

Aug 11, 14:50 UTC
Investigating - We are investigating reports of degraded performance for API Requests

View →
InngestAug 10, 2026 at 22:30 UTC

Delayed function execution for 15 minutes

Status: Resolved

We observed temporary function execution delays for a period of 15 minutes as we performed emergency maintenance to manage disk capacity to several nodes.

Affected components
  • Inngest Dashboard (Operational)
  • Function execution (Operational)
  • API (REST and GraphQL) (Operational)
  • Observability (Operational)
  • Event API (Operational)
View →
GitHubresolvedAug 10, 2026 at 20:27 UTC

Disruption with Copilot for access to some models

Aug 10, 21:50 UTC
Resolved - On August 10, 2026, between 19:48 UTC and 20:49 UTC, GitHub Copilot users saw an incomplete list of available models. During this window, the service could return as few as one model instead of the full catalog. Requests that tried to use a model missing from that shortened list failed with a "model not found" error. Copilot requests that used an available model were not affected. This did not affect customers on data-residency (Proxima) environments.

The issue was caused by a change to how model data was published, which our systems could not read back correctly and fell back to a limited default list.

We mitigated the incident by 20:49 UTC and deployed a fix to prevent immediate recurrence by 21:50 UTC. We are adding validation and retry safeguards so that model data is verified before it is served.

We apologize for the disruption.

Aug 10, 21:50 UTC
Update - We have deployed and validated the fix to prevent immediate reoccurrence. We will be performing additional work to limit these kinds of failures in the future.

Aug 10, 21:19 UTC
Update - The issue has been mitigated across all affected environments. We are currently deploying on a fix to prevent reoccurrence. We will provide another update once the fix has been deployed.

Aug 10, 20:49 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Aug 10, 20:39 UTC
Monitoring - We are currently investigating reports of some Copilot users experiencing issues accessing certain models. Affected users may see errors or degraded functionality when attempting to use specific models. We are actively working on a fix and will provide updates as we have more information.

Aug 10, 20:27 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 10, 2026 at 18:02 UTC

Disruption with creation of fine grained personal access tokens

Aug 10, 18:46 UTC
Resolved - On August 10, 2026, between 17:16 and 18:21 UTC, users were unable to create new fine-grained personal access tokens (FG PAT) through the GitHub website. When a user submitted the FG PAT creation form, they were returned to the FG PAT list without an error message and no FG PAT was created. Creating classic personal access tokens, as well as editing or deleting existing FG PAT were not affected.

The cause was a change to how the website loads certain front-end JavaScript that was enabled for all users at 17:15 UTC; the change interacted with an issue in the token creation form's confirmation step that prevented it from running, so the final submission that actually creates the token never completed. Because the page still loaded and the server returned a normal response, the failure produced no error message. GitHub mitigated the incident by disabling the change at 18:21 UTC, at which point token creation recovered immediately, and the incident was resolved at 18:46 UTC.

To reduce the chance of recurrence, GitHub is adding monitoring and alerting for anomalies in the FG PAT creation success rate and is removing the issue in the FG PAT creation form that prevented the confirmation step from running. GitHub is also adding automated detection of the issue so other areas of the GitHub front end do not repeat the problem.

Aug 10, 18:22 UTC
Update - We identified the source of the issue affecting creation of fine-grained personal access tokens and have applied a mitigation. Users should now be able to create new fine-grained tokens successfully. We are continuing to monitor to confirm full recovery.

Aug 10, 18:21 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Aug 10, 18:09 UTC
Update - We are investigating reports of users being unable to create fine-grained Personal Access Tokens. Attempting to create a new token redirects the user back to the token overview page without an error message, but the token was not created.

Aug 10, 18:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendAug 9, 2026 at 19:22 UTC

Delayed contact webhook events

Status: Resolved

We have resolved the underlying issue and service has been resumed.
View →
ResendAug 8, 2026 at 22:29 UTC

Increased API and SMTP Error Rates

Status: Resolved

Our team has mitigated the issue and this incident is fully resolved. All services are operational.

Affected components
  • SMTP (Operational)
  • Batch Emails (Operational)
  • Webhooks (Operational)
  • Single Email (Operational)
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • General API (Operational)
  • Website (Operational)
  • Automations (Operational)
  • Email Events (Operational)
View →
ResendAug 7, 2026 at 15:43 UTC

Delayed contact related webhook events

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Webhooks (Operational)
View →
VercelresolvedAug 6, 2026 at 16:37 UTC

Delays delivering log drains

Aug 6, 18:29 UTC
Resolved - This incident has been resolved. Log drain delivery has recovered and remaining delayed deliveries have been processed.

Aug 6, 18:19 UTC
Monitoring - Log drain delivery has recovered. We are monitoring while remaining delayed deliveries are processed.

Aug 6, 17:48 UTC
Update - We are beginning to see recovery and continue to work on the issue. We will provide updates as they become available.

Aug 6, 17:06 UTC
Identified - We've identified the issue and are working on a fix. We will provide updates as they become available.

Aug 6, 16:58 UTC
Update - We continue to investigate this issue and will provide updates as they become available.

Aug 6, 16:37 UTC
Investigating - We've identified an issue causing delays in log drain delivery. We are investigating and will provide updates as they become available.

View →
GitHubresolvedAug 6, 2026 at 15:22 UTC

Incident with Actions

Aug 7, 02:04 UTC
Resolved - On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes.

The incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services.

These services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events.

As the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the services responsible for job assignment, runners were getting assigned jobs that were no longer valid and then getting stuck retrying those jobs, preventing them from picking up valid work.

This second stage of impact was mitigated by deploying changes to prevent runners from repeatedly attempting to acquire invalid jobs. These mitigations allowed the accumulated queues to drain and Actions to recover to normal operation.

Some Actions Runner Controller (ARC) runners remained stuck after the incident. A mitigation deployed during the incident inadvertently affected these runners, causing some to remain offline until they were manually recovered. We subsequently rolled back the change and are adding automatic recovery in upcoming Runner and ARC releases.

Some jobs created during the incident were also left stuck unable to be retried or canceled. CLI and UI solutions for customers to address these were shared at https://github.com/orgs/community/discussions/204152#discussioncomment-17946043.

To prevent recurrence, we are making improvements to deployment and capacity safeguards for the affected services, strengthening monitoring for the conditions that preceded the incident, improving the resiliency and recovery of queued work and runner assignment, and adding automatic recovery for self-hosted runners affected by similar failure conditions. We are also making additional improvements to reduce the risk of cascading failures and accelerate recovery during large-scale Actions disruptions.

Aug 7, 02:03 UTC
Update - During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners.

The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future.

Some workflow-triggering events, including push and pull request events, were not processed during the incident and cannot be replayed automatically. Customers may need to repeat the triggering action by pushing a new commit, updating the pull request, or manually re-running the workflow where applicable.

Aug 7, 00:59 UTC
Update - We’re investigating reports that some Actions Runner Controller runners are taking longer than expected to recover. We’ll provide an update as our investigation progresses.

Aug 7, 00:06 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Aug 7, 00:05 UTC
Update - The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.

Aug 7, 00:01 UTC
Update - System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.

Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.

We are monitoring all affected services for sustained recovery and will provide another update shortly.

Aug 7, 00:01 UTC
Update -
System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.

Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.

We are monitoring all affected services for sustained recovery and will provide another update shortly.

Aug 6, 23:13 UTC
Update - We have deployed fixes that address runners being assigned invalid jobs and are taking additional steps to clear the backlog of affected jobs. Job completion rates for running workflows have improved significantly, with success rates now at 99%. Global queues for hosted runner assignment are nearly burned down and concurrency queues for customers are being processed. Another change was deployed to accelerate processing the backlog of job requests.

We are gradually restoring throughput for webhook-triggered Actions workflows and monitoring system stability. We have deployed a fix for self-hosted runners that were not picking up jobs and are enabling it incrementally.

GitHub Pages, Copilot code review, and Copilot coding agent may still experience intermittent failures or delays. Migrations using GitHub Enterprise Importer remain paused.

We continue to monitor recovery across all affected services and will provide another update as conditions improve.

Aug 6, 22:18 UTC
Update - We continue to make progress on the issue affecting GitHub Actions. We have deployed a fix that addresses runners being assigned jobs that are no longer valid, and are seeing improvement in job completion rates. For workflow runs that are starting, success rates have increased significantly and are now at 97%. Standard and larger runners are now draining queued work. A change is also in progress to mitigate issues with existing self-hosted runners that are not picking up jobs.

Webhook triggers remain throttled to support recovery. Many push and pull request events are not yet triggering new workflow runs, and we are working to safely restore full throughput.

GitHub Pages, Copilot code review, and Copilot coding agent may still experience failures or delays. Migrations using GitHub Enterprise Importer remain paused.

We are continuing to monitor recovery and will provide another update as conditions improve.

Aug 6, 21:30 UTC
Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers remain throttled to aid recovery, so many push and pull request events are not triggering new workflow runs.

We identified runners being assigned jobs that are no longer valid and are deploying a change to address this issue. Both GitHub-hosted and self-hosted runners are affected.

Copilot code review, Copilot coding agent, and GitHub Pages may experience failures or delays. Migrations using GitHub Enterprise Importer have been paused to support mitigation efforts.

Aug 6, 20:34 UTC
Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.

We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.

Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.

Aug 6, 19:43 UTC
Update - We are continuing to work on an issue affecting GitHub Actions.

Capacity remains constrained and jobs may still be delayed or fail while it recovers gradually. Customers using self-hosted runners may see errors or rate limiting when runners register. 

Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed.

Our engineers remain actively engaged.

Aug 6, 18:46 UTC
Update - We are continuing to work on an issue affecting multiple GitHub services.

Workflow runs are still failing, and jobs may remain queued for an extended period before starting or may time out. Jobs using GitHub-hosted runners are particularly affected while capacity is constrained.

Customers using self-hosted runners may see errors or rate limiting when runners register.

Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed.

Recovery is taking longer than we expected, and engineers remain actively engaged.

Aug 6, 18:11 UTC
Update - We are continuing to work on an issue affecting multiple GitHub services.

Workflow runs are still failing or delayed in starting, and some queued jobs may time out.

Customers using self-hosted runners may see errors or rate limiting when runners register.

Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer may also be affected.

Webhook deliveries may be delayed.

Engineers have applied further mitigations and are continuing to work towards full recovery.

Aug 6, 17:40 UTC
Update - We are continuing to work on an issue affecting multiple GitHub services.

Workflow runs are failing or delayed in starting, and some queued jobs may time out.

Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer might also affected.

Webhook deliveries may be delayed.

Engineers have applied a number of mitigations and are rolling out a further fix across all affected systems now.

Aug 6, 17:02 UTC
Update - We are continuing to work on the issue affecting GitHub Actions.

Workflow runs are still failing or delayed in starting, and some queued jobs may time out.

Some requests to the Actions API are returning errors. Customers running migrations with GitHub Enterprise Importer may see failures.

Our engineers have applied several mitigations and are rolling out a further fix now.

Aug 6, 16:33 UTC
Update - Actions and Pages are experiencing degraded availability. We are continuing to investigate.

Aug 6, 16:27 UTC
Update - We are continuing to work on the issue affecting GitHub Actions.

Some workflow runs are still delayed or failing to complete, and some requests to the Actions API are returning errors.

Customers running migrations with GitHub Enterprise Importer may also see failures.

Engineers are actively working towards full recovery.

Aug 6, 16:27 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Aug 6, 16:19 UTC
Update - Pages is operating normally.

Aug 6, 15:53 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Aug 6, 15:45 UTC
Update - We are investigating errors affecting GitHub Actions. Some workflow runs are failing to start or failing partway through, and some requests to the Actions REST API are returning errors.

Some customers may also see unexpected rate limiting in their workflows.

Engineers have identified the source of the disruption and are actively working on a mitigation

Aug 6, 15:41 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Aug 6, 15:22 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
GitHubresolvedAug 6, 2026 at 15:03 UTC

Incident with Pages - Deployment Lag

Aug 6, 16:22 UTC
Resolved - On August 6, 2026, at 07:00 UTC, a configuration change inadvertently reduced the capacity of the service that processes GitHub Pages deployments. As traffic increased over the following hours, latency in the deployment pipeline progressively increased.

At 12:09 UTC, latency crossed the alerting threshold and the team began investigating. We reverted the invalid configuration and applied additional mitigations, including reducing status deployment processing to lower the load on our Redis cluster. Latency returned to normal levels at 15:40 UTC.

Customer impact occurred from 11:34 to 15:32 UTC. During this period, we failed to process approximately 128,000 deployments.

We have updated our alerts to detect elevated processing latency sooner and to notify us immediately when latency causes deployment processing failures. We've confirmed this incident was not fully captured by our availability metrics. In the coming days, we'll update how GitHub Pages availability is measured so incidents like this are accurately reflected going forward.

Aug 6, 15:50 UTC
Monitoring - The degradation affecting Pages has been mitigated. We are monitoring to ensure stability.

Aug 6, 15:03 UTC
Investigating - We are investigating reports of degraded performance for Pages

View →
GitHubresolvedAug 5, 2026 at 11:38 UTC

Some Copilot Cloud Agent jobs not starting

Aug 5, 13:00 UTC
Resolved - On August 5, 2026, between 11:02 and 11:54 UTC, the GitHub Copilot cloud agent service was degraded and new cloud agent jobs were delayed from starting. During this period 100% of newly submitted agent jobs were affected. The incident was limited to delay of cloud agent jobs. No jobs were lost and the queued backlog was processed by 13:00 UTC. This was due to an internal rate limit used to protect service availability that was enabled more broadly than intended delaying more traffic than expected.

The service recovered when the rate limit window expired. We then tuned the control so it no longer affected unrelated coding agent traffic.

We are working to improve the control's scoping and our monitoring and alerting to reduce our time to detection and mitigation of similar issues in the future.

Aug 5, 12:10 UTC
Update - Copilot cloud agent jobs have recovered and the backlog of delayed jobs is being processed.

Aug 5, 12:01 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Aug 5, 11:38 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedAug 3, 2026 at 09:54 UTC

Incident with Copilot

Aug 3, 11:25 UTC
Resolved - On 2026-08-03, between 06:52 and 11:25 UTC, some GitHub Copilot users experienced errors when using chat and agent features. Requests to list the available models failed, and because every chat or agent interaction begins by retrieving the list of models, affected users saw their requests fail. On average about 3% of these model-listing requests failed during the incident (roughly 97% succeeded), but failures were significantly higher during peak-traffic periods, at times approaching 100% for the affected internal lookups. Approximately 4,066 users were affected in a single 60-minute window, concentrated among IDE-based clients. The underlying AI models themselves remained healthy throughout.

The incident was caused by an increase in how often clients requested the model list, which pushed an internal user-authorization lookup past a rate limit; the rate-limited responses were surfaced to users as errors. We mitigated the impact by increasing how long Copilot caches that authorization lookup, which reduced load on the internal service, and we have additional capacity and rate-limit changes in progress. To prevent recurrence we are improving monitoring for this class of failure, adjusting cache and rate-limit settings, and coordinating with client teams on request patterns.

Aug 3, 11:19 UTC
Monitoring - The degradation affecting Copilot has been mitigated. We are monitoring to ensure stability.

Aug 3, 10:35 UTC
Update - We are still seeing intermittent errors with Copilot, and are continuing to investigate and consider mitigations.

Aug 3, 09:54 UTC
Update - We are experiencing degraded availability for chat & agent models in Copilot. Multiple models are impacted and customers may experience requests failing. We are investigating and will provide an update as soon as possible.

Aug 3, 09:53 UTC
Investigating - We are investigating reports of degraded performance for Copilot

View →
GitHubresolvedAug 1, 2026 at 18:03 UTC

Incident with Copilot AI Model Providers

Aug 1, 18:44 UTC
Resolved - On August 1, 2026, between 17:47 UTC and 18:20 UTC, users of the Fable 5 model in GitHub Copilot experienced increased request failures and latency. The average failure rate across all Copilot requests was 0.007%, while failures for Fable 5 peaked at 5.6%. Other models remained available. This was caused by degradation of an upstream model provider.

The affected endpoint recovered, and we monitored the service until error rates and latency returned to normal levels. We are working to add endpoint redundancy to mitigate similar provider issues in the future.

Aug 1, 18:23 UTC
Update - The issues with our upstream model provider have been resolved, and Fable 5 is once again available in Copilot products and IDE surfaces.

We will continue monitoring to ensure stability, but mitigation is complete.

Aug 1, 18:20 UTC
Monitoring - The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

Aug 1, 18:20 UTC
Update - We are experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Aug 1, 18:03 UTC
Update - We are seeing increased error rates from specific upstream AI Model Providers

Aug 1, 18:03 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
GitHubresolvedAug 1, 2026 at 11:16 UTC

Degraded availability GPT 5.6 Luna

Aug 1, 12:30 UTC
Resolved - On August 1st, 2026, the GPT-5.6 Luna model in GitHub Copilot experienced degraded availability in intermittent time intervals between ~08:05 UTC and ~16:30 UTC. Specifically the timeframes observed were 10:00-10:20 UTC, 10:45-11:50 UTC, 13:00-14:25 UTC, and 16:00-16:30 UTC. During this time, requests to GPT-5.6 Luna in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for GPT-5.6 Luna fully recovered once the provider resolved their outage at 16:30 UTC.

Aug 1, 12:29 UTC
Update - The issues with our upstream model provider have been resolved, and GPT-5.6 Luna is once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability, but mitigation is complete.

Aug 1, 12:13 UTC
Update - We keep working with our upstream model provider, and are observing recovery. We continue monitoring to ensure stability.

Aug 1, 11:20 UTC
Update - We are experiencing degraded availability for the GPT-5.6 Luna model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot

Aug 1, 11:16 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
ResendAug 1, 2026 at 00:00 UTC

Contact webhook delivery backlog

Status: Resolved

The backlog is gone, and all delayed contact webhooks have been delivered. Delivery latency has now recovered. The new infrastructure configurations prevent this same issue from happening again.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • SMTP (Operational)
  • General API (Operational)
  • Automations (Operational)
View →
ResendJul 30, 2026 at 21:25 UTC

Increased failure rate in Domain verifications

Status: Resolved

We have resolved the underlying issue, and service has resumed. Domain verification has been restored, and affected verifications were re-triggered.

Affected components
  • Dashboard (Operational)
  • General API (Operational)
View →
ResendJul 30, 2026 at 16:23 UTC

Test domain resend.dev failing test emails

Status: Resolved

We have resolved the underlying issue, and sending test emails from resend.dev: http://resend.dev has resumed.

Affected components
  • Automations (Operational)
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
View →
GitHubresolvedJul 30, 2026 at 09:07 UTC

Copilot model Claude Fable 5 experiencing elevated errors

Jul 30, 10:12 UTC
Resolved - On July 30, 2026, the Claude Fable 5 model in GitHub Copilot experienced degraded availability for approximately 73 minutes, from 08:33 to 09:46 UTC. During this time, requests to Claude Fable 5 in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for Claude Fable 5 fully recovered once the provider resolved their outage at 09:46 UTC, and we confirmed resolution at 10:12 UTC.

Jul 30, 10:11 UTC
Update - The issues with our upstream model provider have been resolved, and Claude Fable 5 is once again available in Copilot products and IDE surfaces.

We will continue monitoring to ensure stability, but mitigation is complete.

Jul 30, 09:17 UTC
Update - We are experiencing degraded availability for the Claude Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Jul 30, 09:07 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
ResendJul 29, 2026 at 21:14 UTC

Elevated error rates on email sending, API, and dashboard

Status: Resolved

Between 20:15 and 20:16 UTC, a brief database disruption caused elevated error rates for approximately 45 seconds. During this window, requests to the Resend API, SMTP relay, and dashboard may have returned 5xx errors or timed out. The issue resolved automatically at 20:16 UTC and all services have been operating normally since. Any request that failed during this window can be safely retried. Emails that had already been accepted before the disruption were not lost. They were delivered normally. Broadcasts were not affected.

Affected components
  • Batch Emails (Operational)
  • General API (Operational)
  • Single Email (Operational)
  • Dashboard (Operational)
  • SMTP (Operational)
View →
GitHubresolvedJul 29, 2026 at 20:07 UTC

Incident with Copilot AI Model Providers

Jul 29, 21:51 UTC
Resolved - On July 29, 2026, between 19:45 UTC and 21:51 UTC, users of the Fable 5 model in GitHub Copilot experienced increased request failures and latency. The average failure rate across all Copilot requests was 0.006%, while failures for Fable 5 peaked at 21%. Other models remained available. This was caused by degradation of an upstream model provider.

The affected endpoint recovered, and we monitored the service until error rates and latency returned to normal levels. We are working to add endpoint redundancy to mitigate similar provider issues in the future.

Jul 29, 21:51 UTC
Update - The external ai model provider has resolved the issues, and we have verified Copilot's traffic is fully recovered.

Jul 29, 21:08 UTC
Update - The external AI model provider is continuing to investigate.

Jul 29, 20:38 UTC
Update - The external AI model provider has identified the issue and is working to resolve.

Jul 29, 20:18 UTC
Update - We are investigating increased error rates affecting GitHub Copilot requests to external AI model providers. Some users may experience failures or degraded performance when using Copilot features.

Jul 29, 20:07 UTC
Update - We are seeing increased error rates with requests to specific model providers.

Jul 29, 20:07 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
ResendJul 29, 2026 at 17:39 UTC

Increased returned error rates

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Single Email (Operational)
View →
GitHubresolvedJul 29, 2026 at 15:26 UTC

Incident with Actions

Jul 29, 16:00 UTC
Resolved - On July 29, 2026, from 14:51 UTC to 15:28 UTC, GitHub Actions experienced elevated REST API request timeouts and errors, failures registering runners, and delayed workflow run starts for customers whose traffic was served by a single infrastructure site. This was caused by an under-provisioned internal Actions service in that site: under increased load its instances ran out of memory and became unresponsive, and because Actions API requests wait synchronously on that service, requests routed through the affected site stalled and timed out. During the incident, approximately 2% of workflows were delayed. Requests served by other sites remained unaffected. Both standard and larger hosted runners routed through the affected site could see delayed job starts.

The issue was mitigated by scaling out the runner-administration service in the affected site and increasing the replica count, which restored API availability and returned workflow run starts to normal. We are working to add horizontal autoscaling, memory-saturation alerting, and scaling-forecast monitoring for this service, along with responder playbooks, to reduce the likelihood of similar issues in the future.

Jul 29, 15:40 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 29, 15:34 UTC
Update - We are investigating an issue affecting GitHub Actions. Some customers may experience timeouts or failures with runner registration and workflow runs may be delayed during startup. Our team is actively working to mitigate the impact by scaling capacity across additional infrastructure.

Jul 29, 15:26 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
GitHubresolvedJul 27, 2026 at 19:31 UTC

Test Incident Posted in Error – No Customer Impact (Will Be Removed)

Jul 27, 19:37 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Jul 27, 19:31 UTC
Update - This is just a test of our systems. Please ignore.

Jul 27, 19:31 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJul 27, 2026 at 03:53 UTC

Incident with GraphQL API Requests

Jul 27, 04:09 UTC
Resolved - On July 26, 2026 at 21:34 UTC we began seeing intermittent errors on the GitHub GraphQL API. A subset of GraphQL API requests returned HTTP 502 errors in short bursts. During the impact window an average of 0.09% of GraphQL API requests in the affected region failed, with a peak of 0.50% of requests failing during the worst two-minute period at 03:02 UTC on July 27. Requests that failed generally succeeded when retried, and no data was lost or altered. Other GitHub services were not affected.

The errors were traced to a single group of servers handling a share of GraphQL API traffic. Application processes on that group intermittently closed connections before completing responses. Impact ended at 03:52 UTC on July 27 when those processes were replaced, and we resolved the incident at 04:09 UTC on July 27 after confirming error rates had returned to normal.

We are still investigating why those processes closed connections, and that work is being carried out by the team that owns the underlying compute platform. In the meantime we are adding detection and automated mitigation for when a single group of servers behaves differently from its peers.

Jul 27, 04:09 UTC
Monitoring - The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

Jul 27, 03:53 UTC
Investigating - We are investigating reports of degraded performance for API Requests

View →
VercelresolvedJul 25, 2026 at 18:57 UTC

Vercel Community temporarily unavailable

Jul 26, 17:46 UTC
Resolved - Maintenance on Vercel Community has concluded. We thank you for your patience, and apologize for any inconvenience.

Jul 25, 18:57 UTC
Investigating - The Vercel Community forums are currently undergoing maintenance. We are working to restore access as soon as possible.

View →
GitHubresolvedJul 25, 2026 at 12:34 UTC

Actions run failures and delays

Jul 25, 13:13 UTC
Resolved - Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/s65j9gslmfm8

Jul 25, 13:12 UTC
Update - We have seen recovery in GitHub Actions performance following our earlier mitigation. Workflow runs are processing normally, though jobs queued before 12:40 UTC may still experience failures and will need to be retried.

Jul 25, 12:59 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 25, 12:58 UTC
Update - We have applied a mitigation for the infrastructure issue affecting GitHub Actions. Workflow run failures and delays are improving but not yet fully resolved. Our engineering team continues to work on restoring full functionality across all affected infrastructure.

Jul 25, 12:34 UTC
Update - We are experiencing issues with GitHub Actions that are causing workflow run failures and delays for some users. Our engineering team is actively investigating the infrastructure issue and working to restore full functionality.

Jul 25, 12:31 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
GitHubresolvedJul 25, 2026 at 09:42 UTC

Several GPT models degraded

Jul 25, 10:11 UTC
Resolved - On July 25, 2026, between 09:07 and 10:04 UTC, the GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna models experienced degraded availability in GitHub Copilot products and IDE surfaces. Requests to these models had an average failure rate of 5.6%. Other Copilot models remained available as alternatives.

The degradation was caused by an issue with an upstream model provider. Success rates returned to normal after the upstream issue was mitigated, and we continued monitoring before resolving the incident. We are working on improving the automated failover for the affected models to prevent similar incidents in the future.

Jul 25, 10:04 UTC
Monitoring - The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

Jul 25, 09:48 UTC
Update - We are experiencing degraded availability for the GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna models in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Jul 25, 09:42 UTC
Investigating - We are investigating reports of degraded availability for Copilot AI Model Providers

View →
GitHubresolvedJul 25, 2026 at 08:59 UTC

Incident with Actions

Jul 25, 09:25 UTC
Resolved - On July 25, 2026, GitHub Actions experienced two related periods of degradation that caused some workflow runs to be delayed by more than 5 minutes or end with infrastructure failures.

First period (08:45 – 09:13 UTC): During planned maintenance on a critical-path Redis cluster for Actions, one participating region was left in a degraded state. Separately, an independent capacity operation temporarily removed another region from the cluster and redirected its traffic to the degraded region. This created cross-region inconsistencies in job-assignment state, causing workflow runs to be delayed, exhaust retries, or fail outright. At peak, about 7% of runs were delayed by more than 5 minutes, and 25% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 09:13 UTC by returning traffic to its normal distribution.

Second period (12:08 – 12:48 UTC): As part of mitigating the first incident, traffic was returned to the regional instance that was still undergoing its capacity increase. Multiple Redis nodes in the scaling region experienced failures, increasing traffic to healthy nodes and causing connection limits to be reached on many nodes. At peak, 30% of runs were delayed by more than 5 minutes, and 60% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 12:48 UTC by redirecting workflow traffic away from the scaling region.

We are adding stronger regional health and capacity checks before maintenance and requiring a stable observation period before restoring traffic. We are also improving automated connection resiliency, and partnering with our platform dependency to automatically detect and remediate unhealthy cluster members and shard imbalance. More generally, we already had work underway to improve the resiliency and scale of this piece of Actions infrastructure.

Jul 25, 09:20 UTC
Update - We identified an issue causing delays in GitHub Actions run starts. Some users may have experienced longer than expected wait times when triggering workflow runs. We have applied mitigations and have recovered. Our team continues to monitor and investigate the root cause.

Jul 25, 09:13 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 25, 08:59 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
GitHubresolvedJul 24, 2026 at 19:37 UTC

Incident with Pull Requests

Jul 24, 20:23 UTC
Resolved - Between July 24, 19:17 UTC and July 24, 20:02 UTC, users were unable to create pull requests due to a database schema change. In total, 113,930 pull request creation attempts were impacted across 50,904 users, with an average error rate of 1.75% and a maximum error rate of 2.25% for all requests to Pull Requests service. Existing pull requests and other GitHub functionality were not affected. The issue was resolved by reverting the change to the affected database, upon which pull request creation immediately resumed.

The root cause was related to a backfill workflow into the Vitess keyspace hosting Pull Request data. The backfill Vitess command encountered errors and increased VReplication lag, and the workflow was canceled at 19:17 UTC. The cancellation executed a misunderstood Vitess codepath that dropped the backing table to the target keyspace, leaving a non-existent reference that resulted in errors creating Pull Requests. The mitigation was executing a command to drop the vschema reference to the dropped table, allowing Pull Request creation to resume.

We are adding stronger pre-flight validation to our tooling to prevent similar issues and expanding lower-environment support to provide better test coverage end-to-end before promoting them to production. We're also fixing our backfill migration tooling to protect from this specific codepath.

Jul 24, 20:02 UTC
Monitoring - The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

Jul 24, 19:59 UTC
Update - We have applied a mitigation and are monitoring for recovery

Jul 24, 19:43 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

Jul 24, 19:40 UTC
Update - We are investigating errors creating pull requests

Jul 24, 19:37 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
GitHubresolvedJul 24, 2026 at 16:19 UTC

Disruption with some GitHub services

Jul 24, 17:36 UTC
Resolved - On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ.

Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors:

- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts.
- 27% of GitHub issues interactions saw slow requests or timeouts.
- 4% of GitHub Copilot requests experienced errors, though most automatically retry.
- 4% of git push operations saw impacts during the affected window.
- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases.

We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36.

This incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss.

Jul 24, 17:24 UTC
Update - We are seeing recovery across all services

Jul 24, 17:16 UTC
Update - The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability.

Jul 24, 16:41 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Jul 24, 16:40 UTC
Update - We have applied a mitigation and are monitoring for recovery

Jul 24, 16:28 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jul 24, 16:27 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Jul 24, 16:26 UTC
Update - Copilot is experiencing degraded performance. We are continuing to investigate.

Jul 24, 16:22 UTC
Update - We are investigating timeouts to some GitHub services

Jul 24, 16:20 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Jul 24, 16:19 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Jul 24, 16:17 UTC
Investigating - We are investigating reports of degraded performance for API Requests and Issues

View →
GitHubresolvedJul 24, 2026 at 11:00 UTC

Incident With Blocked GitHub.com Traffic

Jul 24, 11:00 UTC
Resolved - Between July 23, 2026 at 18:45 UTC and July 24, 2026 at 11:19 UTC, an abuse mitigation update caused some legitimate customers whose traffic was routed through our Central Europe and South America edge locations to be incorrectly blocked from GitHub.com. We estimate that approximately 0.25% of GitHub.com requests were affected during this period.

This was caused by an abuse mitigation configuration that incorrectly classified legitimate traffic. We mitigated the incident by reverting the update. We are adding validation and safeguards to prevent similar incorrect blocking in the future.

View →
VercelresolvedJul 23, 2026 at 19:00 UTC

Telemetry data loss for Drains

Jul 23, 19:00 UTC
Resolved - Between 19:12 and 19:18 UTC on July 23, 2026, telemetry forwarded through Drains was not delivered. Traces and events forwarded via Drains during this window were dropped and are unrecoverable. Logs forwarded via Drains during this window were also not delivered, but remain accessible in the Logs UI dashboard, where they can also be exported.

We have deployed a fix and are adding additional monitoring to detect and prevent this failure mode from happening in the future.

We know how much you rely on this data, and we sincerely apologize for the disruption.

View →
InngestJul 23, 2026 at 15:24 UTC

Function processing delays

Status: Resolved

The incident is now resolved and the system is full operational. Delays began around 14:03 UTC caused by an issue publishing to a new Kafka topic. Recovery began around 14:40 UTC as the team was able to isolate this topic. The consuming services which schedules new function runs was scaled out and began to consume the backlog. No events were dropped during this incident, all functions were scheduled, but with delays during this time window. No action is needed for manual recovery.

Affected components
  • Function execution (Operational)
View →
GitHubresolvedJul 23, 2026 at 07:53 UTC

Latency issues across a number of services

Jul 23, 09:39 UTC
Resolved - On July 23, 2026, between 07:08 and 09:39 UTC, several services experienced delays: 8% of actions workflow runs experienced an average run start delay of 10 minutes, 5% of webhook deliveries exceeded SLO, and code scanning, repos, notifications, issues and pull requests experienced increased latency over the life of the incident.

The root cause of the incident was a node of our background job processing system which did not recover after entering scheduled host maintenance. The incident was mitigated by identifying the problematic shard and restoring its correct state, after which queue backlogs drained and services recovered.

To speed mitigation, we have added monitors for nodes in this unhealthy state after maintenance operations. To prevent future recurrence, we are adapting our lifecycle automation to verify host rejoin after a scheduled reboot.

Jul 23, 09:39 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jul 23, 09:35 UTC
Update - Webhooks is operating normally.

Jul 23, 09:27 UTC
Update - The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

Jul 23, 09:22 UTC
Update - We identified the source of latency affecting multiple services and applied a fix. Issues and Actions are recovering, and remaining affected services are seeing improvement as processing backlogs clear. We are actively monitoring recovery across all services.

Jul 23, 09:19 UTC
Update - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 23, 09:18 UTC
Update - The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

Jul 23, 08:34 UTC
Update - We're currently investigating latency across multiple services. This can show as Actions jobs taking longer to start, Issues search serving stale results, and other listed services being similarly impacted.

Jul 23, 08:25 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Jul 23, 07:53 UTC
Investigating - We are investigating reports of degraded availability for Actions, Issues and Webhooks

View →
VercelresolvedJul 23, 2026 at 07:33 UTC

Timeouts loading charts and observability data in Dashboard

Jul 23, 08:18 UTC
Resolved - This incident has been resolved.

Jul 23, 07:55 UTC
Monitoring - The team identified the source of the issue and a fix has been implemented. We are starting to see recovery. Charts and observability data should be loading again for Firewall, Observability, Sandboxes, Speed Insights, Web Analytics, and Workflows. We continue to monitor.

Jul 23, 07:33 UTC
Investigating - We are currently investigating timeouts loading observability data for Firewall, Observability, Sandboxes, Speed Insights, Web Analytics, and Workflows. This impacts charts in the Dashboard.

View →
VercelresolvedJul 23, 2026 at 06:47 UTC

Dashboard authentication errors

Jul 23, 07:41 UTC
Resolved - This incident has been resolved.

Jul 23, 07:08 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jul 23, 06:47 UTC
Investigating - We are investigating an issue causing the dashboard to be inaccessible for some users.

View →
GitHubresolvedJul 22, 2026 at 20:43 UTC

Disruption with actions hosted runners

Jul 22, 22:09 UTC
Resolved - On July 22, 2026, between 19:36 UTC and 22:04 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 15% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 1% failed to start.

At 21:49 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents.

Jul 22, 22:01 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 22, 22:01 UTC
Update - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 22, 20:47 UTC
Update - Approximately 3% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A small portion of these runs may fail after extended delays. We have identified the cause and are working on a mitigation.

Jul 22, 20:43 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
ResendJul 22, 2026 at 19:51 UTC

Errors creating domains and API keys

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Website (Operational)
  • Automations (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Webhooks (Operational)
View →
ResendJul 21, 2026 at 19:39 UTC

Delivery Delayed Events to Microsoft Hosted Domains

Status: Resolved

Delivery to Microsoft has returned to normal. Thank you for your patience as we worked with them regarding this.
View →
GitHubresolvedJul 21, 2026 at 10:31 UTC

Some SSH connections using deploy keys are failing

Jul 21, 11:57 UTC
Resolved - On July 21, 2026, between 07:41 UTC and 11:57 UTC, the SSH Authentication service was degraded and some SSH connections failed to authenticate. On average, 12.2% of SSH authentication requests failed, peaking at 15.7%. Both user RSA keys and deploy keys were impacted. This was due to a change in how our SSH service handled one public-key authentication method that caused the affected authentication attempts to be rejected as invalid.

We mitigated the incident by reverting the change, after which SSH authentication returned to normal.

We are working to expand our automated test coverage for our SSH public-key authentication flows to catch more edge cases and to improve observability and alerting on SSH authentication failures, to reduce our time to detection and mitigation of issues like this one in the future.

Jul 21, 11:46 UTC
Update - We have identified a recent code change as a potential cause of the SSH authentication failures affecting deploy key connections. Our engineering team is rolling back this change. Customers using deploy keys for SSH access to repositories may continue to experience intermittent connection failures until the fix is deployed.

Jul 21, 11:04 UTC
Update - We are investigating reports of intermittent SSH authentication failures affecting connections that use deploy keys. Customers may experience failed SSH connections when interacting with repositories via deploy keys. Our engineering team is actively investigating the root cause and working toward resolution.

Jul 21, 10:31 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJul 20, 2026 at 16:04 UTC

Disruption with GPT 5.3 Codex

Jul 20, 18:37 UTC
Resolved - Between 06:39 and 18:11 UTC on July 20, 2026, the Copilot service experienced a degradation of the GPT 5.3 model due to an issue with our upstream provider. The upstream model provider returned intermittent errors for GPT 5.3 Codex requests, which caused some responses to fail. Auto mode requests that had selected GPT 5.3 Codex were also impacted. On average about 2% of GPT 5.3 Codex requests failed during this window. Copilot automatically routed eligible traffic away from the impacted provider to reduce customer impact. No other models were impacted.

We worked with the upstream provider throughout the incident and confirmed sustained recovery before resolving.

Jul 20, 18:25 UTC
Update - We are seeing signs of sustained recovery from the upstream provider for GPT 5.3 Codex requests. Requests are succeeding as expected. We are awaiting confirmation from the provider that the issue is fully mitigated before resolving this incident.

Jul 20, 17:21 UTC
Update - The upstream provider continues to return intermittent errors for GPT 5.3 Codex requests. Some users may experience failed or interrupted responses when using this model. Traffic continues to be automatically rerouted to reduce impact, and we are actively working with the upstream provider on a resolution. In the meantime, selecting an alternative model will avoid disruption.

Jul 20, 16:39 UTC
Update - We identified that an upstream provider is returning errors for GPT 5.3 Codex requests, causing some users to experience failed or interrupted responses. Traffic is being automatically rerouted to mitigate impact. We are working with the upstream provider to resolve the underlying issue. We recommend using a different model at this time while we resolve the problem.

Jul 20, 16:04 UTC
Update - We are investigating GPT Codex 5.3 performance.

Jul 20, 16:03 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
GitHubresolvedJul 20, 2026 at 00:25 UTC

Disruption with some GitHub services

Jul 20, 01:46 UTC
Resolved - Between July 19, 2026, at 23:05 UTC and July 20, 2026, at 03:55 UTC, Actions self-hosted and larger runners were unable to connect to GitHub. During this period, Actions jobs were delayed or failed when trying to acquire a runner. Jobs using standard and Mac hosted runners were not affected. Reconnection traffic from affected runners also increased load on GitHub APIs, resulting in 3-4 seconds of additional average request latency and elevated 5xx error rates.

The incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration that disrupted runner connectivity. We restored service by rotating the affected certificate. Recovery began at 02:45 UTC. By 03:55 UTC, queued workflow backlog had been processed and workflow delay rates returned to normal.

To prevent recurrence, we are strengthening certificate renewal automation, adding fallback expiry monitoring and alerting, and improving circuit-breaker protections during runner API disruptions to reduce the risk of cascading impact to other APIs.

Jul 20, 01:45 UTC
Update - Git LFS API success rates have returned to normal. We will continue to monitor the service closely.

Jul 20, 01:15 UTC
Investigating - The Git LFS API issues have been identified as being related to a separate incident where Actions is experiencing degraded availability. We’re consolidating our investigation efforts under that incident.

Jul 20, 00:26 UTC
Update - Some Git LFS operations, and loading files through the API are failing. We are investigating.

Jul 20, 00:25 UTC
Monitoring - Some Git LFS operations, and loading files through the API are failing. We are investigating.

Jul 20, 00:25 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJul 19, 2026 at 23:34 UTC

Incident with GitHub Actions

Jul 20, 04:44 UTC
Resolved - Between July 19, 2026, at 23:05 UTC and July 20, 2026, at 03:55 UTC, Actions self-hosted and larger runners were unable to connect to GitHub. During this period, Actions jobs were delayed or failed when trying to acquire a runner. Jobs using standard and Mac hosted runners were not affected. Reconnection traffic from affected runners also increased load on GitHub APIs, resulting in 3-4 seconds of additional average request latency and elevated 5xx error rates.

The incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration that disrupted runner connectivity. We restored service by rotating the affected certificate. Recovery began at 02:45 UTC. By 03:55 UTC, queued workflow backlog had been processed and workflow delay rates returned to normal.

To prevent recurrence, we are strengthening certificate renewal automation, adding fallback expiry monitoring and alerting, and improving circuit-breaker protections during runner API disruptions to reduce the risk of cascading impact to other APIs.

Jul 20, 04:43 UTC
Update - Actions has fully recovered, and we will continue to monitor the platform to ensure stability.

Jul 20, 03:34 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 20, 03:27 UTC
Update - We’re seeing recovery across all impacted Actions runners. The team is continuing to monitor for global recovery.

Jul 20, 03:03 UTC
Update - API Requests is operating normally.

Jul 20, 02:43 UTC
Update - Issues, API Requests, and Pages have recovered. We are continuing to work on restoring GitHub Actions jobs using self-hosted or larger-hosted runners.

Jul 20, 02:23 UTC
Update - API Requests is experiencing degraded performance. We are continuing to investigate.

Jul 20, 02:20 UTC
Update - Issues is operating normally.

Jul 20, 01:37 UTC
Update - Pages is operating normally.

Jul 20, 01:11 UTC
Update - We continue to work on mitigative efforts to restore Actions workflow runners, and have observed that the extended downtime has started to cause knock-on effects to other services.

A separate incident was opened before we understood they were related. We will continue to post updates on this incident.

Jul 20, 00:52 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Jul 20, 00:49 UTC
Update - Actions and Pages are experiencing degraded performance. We are continuing to investigate.

Jul 20, 00:48 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Jul 20, 00:20 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jul 20, 00:07 UTC
Update - API Requests is experiencing degraded availability. We are continuing to investigate.

Jul 19, 23:57 UTC
Update - We have identified the cause of failures in GitHub Actions and are working to restore service.

Jul 19, 23:37 UTC
Update - We are investigating degraded availability for GitHub Actions on github.com and in GHEC DR stamps. New workflows may delay or fail to start, and ongoing runs may fail. We will provide more information as soon as we can.

Jul 19, 23:34 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
VercelresolvedJul 17, 2026 at 14:13 UTC

Increased invocation failures for Hobby Team functions

Jul 17, 15:19 UTC
Resolved - This incident has been resolved.

A subset of deployments created under the Hobby plan teams during Jul 17, 2026 10:18 - Jul 17, 2026 14:10 UTC were impacted by this incident. New deployments after this window are unaffected.

Jul 17, 14:27 UTC
Monitoring - We have deployed a fix and are continuing to monitor. Function invocations for new Hobby Team deployments should no longer fail. If you are still experiencing function invocation errors, redeploy or rollback to a previous deployment.

Jul 17, 14:19 UTC
Identified - We have identified the source of errors and are implementing a fix.

Jul 17, 14:12 UTC
Investigating - There are elevated rates of invocation failures for Hobby Team functions in new deployments. Existing deployments are unaffected. We are investigating. Mitigation is possible by rolling back to a previous deployment.

View →
VercelresolvedJul 16, 2026 at 23:09 UTC

GitHub-linked deployments and authentication affected

Jul 17, 00:42 UTC
Resolved - This incident has been resolved.

Jul 17, 00:07 UTC
Monitoring - Deployments for projects linked to GitHub repositories are recovering. Some customers may continue to experience delayed build starts or builds stuck in an initializing state while recovery continues. We will provide additional updates as they become available. GitHub login, signup, and account connections are recovered.

Jul 16, 23:39 UTC
Identified - We've identified an issue affecting deployments for projects linked to GitHub repositories, as well as GitHub login, signup, and account connections. CLI deployments and other sign-in methods are unaffected. We will share updates as they become available.

Jul 16, 23:09 UTC
Investigating - We are investigating elevated error rates for users logging in to Vercel using their GitHub account and connecting their GitHub accounts to Vercel. Login with email or other sign-in methods is unaffected.

View →
GitHubresolvedJul 16, 2026 at 22:51 UTC

Degraded REST API Availability

Jul 17, 00:14 UTC
Resolved - From 22:21 UTC - 23:50 UTC on July 16, 2026, the REST API experienced significant degradation. During this period, about 39% of REST API requests failed with HTTP 500 level responses, with the errors peaking at 44.3%.

We identified the issue as an infrastructure change that wrongly marked the majority of API backends in a single region as unhealthy. As a result, requests routed to those backends failed before reaching the application layer.

To prevent this from happening again, we're improving our systems to catch this kind of invalid configuration before it reaches production. We'll also audit the related systems to make them more resilient to future changes, and we're increasing our monitoring sensitivity so we're alerted to problems like this sooner.

Jul 17, 00:14 UTC
Update - As of 23:46 UTC, the REST API service is reachable and responding to requests normally.

Jul 17, 00:00 UTC
Monitoring - The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

Jul 16, 23:29 UTC
Update - We are continuing to investigate an issue causing approximately 35% of REST API requests to fail.  Based on our current understanding, requests are not consistently reaching the application layer, resulting in failed requests returning HTML responses instead of the expected API response format. We are actively investigating the issue and will provide another update as soon as more information is available.

Jul 16, 22:58 UTC
Update - API Requests is experiencing degraded performance. We are continuing to investigate.

Jul 16, 22:58 UTC
Update - We are aware of degraded REST API availability and are investigating

Jul 16, 22:51 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJul 16, 2026 at 21:06 UTC

Claude Fable 5 experiencing degraded performance

Jul 16, 22:04 UTC
Resolved - On July 16, 2026, GitHub Copilot users experienced elevated errors when using Claude Fable 5 from 17:33 UTC until mitigation at 22:04 UTC. The average error rate was 1.4%, with a maximum error rate of 30.85%. The issue was caused by degradation at an upstream model provider; other Copilot models were not significantly affected, and users could avoid the impact by selecting another model or Auto. Service recovered after the provider mitigated the degradation.

Jul 16, 21:06 UTC
Update - We are experiencing degraded availability for the Claude Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Jul 16, 21:05 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
InngestJul 16, 2026 at 19:29 UTC

Metrics Data Delays

Status: Resolved

Metrics in the UI are now restored. Users may notice metrics data continue to be unavailable during the time of the incident, from the timeframe 7/16/26 15:45 UTC until approximately 7/16/26 21:07 UTC

Affected components
  • Inngest Dashboard (Operational)
  • Observability (Operational)
View →
InngestJul 16, 2026 at 12:57 UTC

Latency and Execution Errors

Status: Resolved

Throughput and latency have recovered and the system should be fully operational again. The core issue is related to the part of the system that persists run state. Some operations experienced and increase in errors and retries, causing a backlog in some parts of the system and some failed operations like checkpoints or signals. We are continuing a thorough post-mortem investigation to prevent reoccurrance.

Affected components
  • Function execution (Operational)
View →
ResendJul 16, 2026 at 12:30 UTC

Elevated errors on the dashboard

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • General API (Operational)
  • Website (Operational)
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • Webhooks (Operational)
  • Automations (Operational)
View →
ResendJul 16, 2026 at 11:50 UTC

Domain-related webhooks delayed

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Webhooks (Operational)
View →
GitHubresolvedJul 16, 2026 at 09:13 UTC

Disruption with some GitHub services

Jul 16, 12:20 UTC
Resolved - On July 16, 2026, between 08:50 UTC and 09:50 UTC, the GitHub MCP Server’s web_search tool experienced elevated failures. The average error rate was 42% and peaked at 82% of requests to the tool. Other GitHub MCP Server tools were unaffected. This was caused by degradation at a downstream web search provider.

The incident was mitigated when the downstream provider recovered, after which we confirmed that the tool’s success rate had returned to normal.

We are improving the tool’s resilience and failure handling to reduce the customer impact and duration of similar incidents.

Jul 16, 10:17 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Jul 16, 09:30 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jul 16, 09:13 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedJul 16, 2026 at 08:39 UTC

SAML Single Sign-On (SSO) errors

Jul 16, 10:45 UTC
Resolved - This incident has been resolved.

Jul 16, 10:23 UTC
Monitoring - The root cause of the SAML Single Sign-On (SSO) errors has been identified and a fix has been implemented. We are seeing error rates decrease. Users should be able to login with SAML/SSO or access their SSO-protected teams through the Vercel dashboard and CLI again. We are continuing to monitor.

Jul 16, 08:39 UTC
Investigating - We are investigating SAML Single Sign-On (SSO) errors, leading to inability to login using SAML/SSO or access SSO-protected teams through the Vercel dashboard and CLI.

View →
VercelresolvedJul 16, 2026 at 06:07 UTC

Missing Build CPU Minutes Usage Data

Jul 16, 09:27 UTC
Resolved - The missing Build CPU Minutes data has been backfilled. Users with missing Build CPU Minutes data will be able to see it on Usage pages in the Vercel Dashboard.

Jul 16, 06:07 UTC
Monitoring - We've identified an issue where some users may see missing Build CPU Minutes data on Usage pages in the Vercel Dashboard. The issue has been resolved, and we are backfilling the affected usage data.

View →
VercelresolvedJul 15, 2026 at 23:04 UTC

Errors accessing teams that require 2FA

Jul 15, 23:34 UTC
Resolved - This incident has been resolved.

Jul 15, 23:16 UTC
Monitoring - A fix has been implemented and services have recovered. We are continuing to monitor to ensure that the services remain stable. We will provide additional updates as they become available.

Jul 15, 23:04 UTC
Identified - We have identified an increase in erroneous 2FA (two-factor authentication) challenges and errors when accessing the Vercel dashboard and multiple API services for teams that require 2FA. We will share more information once we have it.

View →
GitHubresolvedJul 14, 2026 at 17:38 UTC

Incident with Webhooks

Jul 14, 18:01 UTC
Resolved - On July 14, 2026, between 15:17 and 15:37 UTC, a rollout to GitHub's internal webhook delivery pipeline caused a subset of webhook delivery records to not be written to our webhook deliveries store after being processed and delivered successfully. Affected deliveries would be missing from the webhook delivery UI and API and won’t be available for redelivery.

The root cause was an uncoordinated rollout: a change to how delivery records are handed off between pipeline components was deployed before the upstream components producing those records were updated to match. While the rollout was in progress, affected records were silently skipped rather than persisted, with no automatic retry. The impact ended as soon as the rollout was completed.

About 2.4M delivery records were skipped (approximately 4% of the 20-minute impact window, 0.04% of a typical 24-hour period). Importantly, 95% of these deliveries reached customer endpoints successfully, only the record of the delivery is missing. Of the ~5% that failed to reach customer endpoints, only ~1.4% (5,463) map to webhooks that retried their deliveries in the past 28 days.

To prevent recurrence, we are making delivery persistence safe-by-default, adding alert on drops in the persisted-delivery rate, and tightening rollout coordination for changes that span multiple components in the pipeline.

Jul 14, 17:55 UTC
Monitoring - The degradation affecting Webhooks has been mitigated. We are monitoring to ensure stability.

Jul 14, 17:55 UTC
Update - Between 15:17 and 15:27 UTC, an ongoing deployment had an unintended side effect where webhook delivery states were not persisted in all cases, even when deliveries were accepted and processed. Customers may notice missing webhook deliveries in the UI, and these deliveries will not be retryable.

Impact resolved automatically once the deployment completed.

Jul 14, 17:38 UTC
Investigating - We are investigating reports of degraded performance for Webhooks

View →
GitHubresolvedJul 14, 2026 at 08:22 UTC

Disruption with some GitHub services

Jul 14, 09:56 UTC
Resolved - On July 14, 2026, the GitHub Codespaces service was degraded during two periods — between 06:00 UTC and 09:56 UTC, and again between 10:54 UTC and 12:53 UTC — and some users experienced intermittent failures or delays when creating new codespaces. Impact was concentrated in a subset of geographic regions. During the first period, the error rate averaged 0.5% and peaked at 4.6% of codespace creation requests. The second period was more pronounced, peaking at approximately 30% of codespace creation requests in the most-affected region before recovery. Both periods were caused by an unexpected surge in codespace creation from an abusive actor that drained the available compute capacity in the affected regions faster than it could be replenished.

We mitigated the impact by identifying and stopping the sources of the excess creation volume, reducing the resources that could be consumed in the affected regions, and rebalancing traffic across regions to restore capacity. Codespace creation success rates returned to normal after each period.

We are working to add automated, low-latency controls to throttle abnormal codespace creation and to strengthen our detection and safeguards, so we can reduce our time to detection and mitigation of issues like this in the future.

Jul 14, 09:51 UTC
Monitoring - The degradation affecting Codespaces has been mitigated. We are monitoring to ensure stability.

Jul 14, 08:22 UTC
Update - Codespaces is experiencing degraded performance. We are continuing to investigate.

Jul 14, 08:21 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
InngestJul 13, 2026 at 19:44 UTC

Issues Affecting Function Triggering from Events

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
GitHubresolvedJul 13, 2026 at 13:32 UTC

Actions runs are experiencing failures to start

Jul 13, 13:53 UTC
Resolved - On July 13, 2026, between 13:11 and 13:53 UTC, some customers experienced failures starting and running GitHub Actions workflows, which also affected Copilot cloud agent sessions and GitHub Pages builds since they depend on Actions. During the peak of the incident, 30% of Actions jobs failed to start and 2% were delayed more than 5 minutes.

The incident was triggered by a configuration change in an internal autoscaling component that contained outdated capacity threshold values. This caused a critical Actions service to scale below its required baseline, reducing capacity for workflow processing. We identified the regression, rolled back the change, and restored service capacity. New workflow executions recovered by 13:39 UTC. Full recovery was reached by 13:53 UTC after the queued backlog was drained.

To prevent recurrence, we have added deployment guardrails to validate that autoscaling inputs are current and to detect drift between planned and live scaling state before autoscaling changes are applied.

Jul 13, 13:39 UTC
Monitoring - The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.

Jul 13, 13:32 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Jul 13, 13:32 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
ResendJul 12, 2026 at 20:29 UTC

Increased latency in CSV imports and automation runs

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Automations (Operational)
View →
VercelresolvedJul 10, 2026 at 21:27 UTC

Increased Build failure rates

Jul 10, 23:28 UTC
Resolved - This incident has been resolved.

Jul 10, 21:22 UTC
Monitoring - We've identified an issue where some customers may experience failed Builds beginning at 21:05 UTC.

A rollback is in progress to mitigate this issue. We will provide additional updates as they become available.

Failed deployments can be retried and should succeed.

View →
VercelresolvedJul 10, 2026 at 19:15 UTC

Erroneous budget notifications

Jul 10, 22:36 UTC
Resolved - We have unpaused all impacted deployments. We will continue to monitor the situation.

Jul 10, 21:32 UTC
Update - We are still working to unpause affected teams that were erroneously paused due to the spend management budget miscalculation.

We will provide additional updates as they become available.

Jul 10, 20:50 UTC
Update - We've compiled a list of deployments that were paused due to this incident and are now unpausing them in batches. If your deployment was paused, you can also navigate to the Vercel dashboard and click the unpause button in the banner.

Jul 10, 19:53 UTC
Identified - We've identified and fixed the underlying issue. Our team is now unpausing deployments that were inadvertently paused as a result.

Jul 10, 19:15 UTC
Investigating - We've identified an issue where some customers may receive erroneous spend management budget notifications due to a spend management budget miscalculation. Affected teams with spend management budgets configured to pause projects may also experience paused deployments. We are currently investigating this issue. We will provide additional updates as they become available.

View →
ResendJul 10, 2026 at 14:14 UTC

Increased latency across API endpoints and email sending

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
View →
VercelresolvedJul 9, 2026 at 13:16 UTC

Delays loading Build Logs

Jul 9, 13:16 UTC
Resolved - The issue affecting Build Logs has been resolved. Between Jul 8 19:05 UTC and Jul 9 13:16 UTC, some Pro customers using standard machine types may have experienced missing Build Logs.

New builds after the impact period are unaffected.

View →
GitHubresolvedJul 9, 2026 at 04:34 UTC

Delays starting Actions runs

Jul 9, 13:52 UTC
Resolved - On July 9, 2026, between 03:29 UTC and 13:39 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 8% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 2% failed to start.

At 13:39 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents.

Jul 9, 13:41 UTC
Update - Actions, Pages builds, Copilot Cloud Agent, and Copilot Code review have all recovered and are mitigated.

We are continuing to monitor to ensure full recovery, and investigating the health of the affected infrastructure.

Jul 9, 13:39 UTC
Monitoring - The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.

Jul 9, 13:17 UTC
Update - We are continuing to monitor slow recovery in Actions and Pages builds as the system works through the high volume of backlog.

Customers may see a small rate of  API and job failures as the system is recovering.

Copilot Cloud Agent and Copilot Code Review also failed to start for approximately 30 minutes during this incident, and we are monitoring recovery.

Pages were accessible throughout the incident.

Jul 9, 13:16 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jul 9, 12:54 UTC
Update - We're seeing Actions and Pages recovery.

For a period of approximate 20 minutes ~96% of GitHub Actions runs on GitHub-hosted runners were failing to start, but has now recovered and we are seeing jobs processing.

GitHub pages builds were also failing during that period, but Pages are still accessible.

We are continuing to monitor for full recovery.

Jul 9, 12:46 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jul 9, 12:36 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

Jul 9, 12:01 UTC
Update - Approximately 30% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A smaller percentage of those are exhausting retries and failing to start.
This has caused some customers to exceed their hosted compute concurrency and experience increased impact.

We are continuing to working on infrastructure mitigations.

Next update in one hour.

Jul 9, 10:15 UTC
Update - Approximately 30% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A smaller percentage of those are exhausting retries and failing to start.

Jul 9, 10:07 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jul 9, 06:01 UTC
Update - We are continuing to work on a mitigation.

Jul 9, 04:51 UTC
Update - Approximately 5% of GitHub Actions runs on GitHub-hosted runners are experiencing run start delays exceeding 5 minutes. A small portion of these runs may fail after extended delays. We have identified the cause and are working on a mitigation.

Jul 9, 04:34 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
VercelresolvedJul 8, 2026 at 18:05 UTC

Delays Processing Builds

Jul 8, 19:13 UTC
Resolved - This incident has been resolved.

Jul 8, 19:13 UTC
Update - The issue affecting build queue processing has been resolved, and build queues have returned to normal. Builds are now processing as expected.

We are continuing our internal root cause analysis.

Jul 8, 19:13 UTC
Monitoring - Build queues have returned to normal levels, and build processing should be restored.

We are monitoring to ensure the service remains stable. We will provide additional updates as they become available.

Jul 8, 18:40 UTC
Update - We are seeing signs of recovery and the affected build queue has started to drain.

Some customers may continue to experience delayed build starts or builds stuck in an initializing state while recovery continues. We will provide additional updates as they become available.

Jul 8, 18:05 UTC
Investigating - We've identified an issue where some customers may experience delays in builds starting and/or builds stuck in an initializing state. We are currently investigating this issue and will provide additional updates as they become available.

View →
InngestJul 8, 2026 at 15:38 UTC

Degraded function execution performance

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
GitHubresolvedJul 7, 2026 at 14:14 UTC

Actions and Codespaces APIs experiencing partial failures

Jul 7, 16:17 UTC
Resolved - On July 7, 2026, between 14:01 UTC and 16:17 UTC the Actions and Codespaces REST APIs were degraded and returned intermittent 500-class errors for a percentage of requests. Error rates peaked at approximately 8% of Actions runner API requests and 13% of Codespaces API requests, though retries were frequently successful. In-progress Actions runs and Codespaces were not impacted and continued successfully. This was due to a recent change that did not deliver the expected performance and, under certain conditions, caused downstream errors.

We mitigated the incident by rolling back the change, after which the affected services recovered.

We are working to improve the resilience of our services to these conditions and to strengthen our monitoring to reduce our time to detection and mitigation of issues like this one in the future.

Jul 7, 16:02 UTC
Monitoring - The degradation affecting Actions and Codespaces has been mitigated. We are monitoring to ensure stability.

Jul 7, 16:00 UTC
Update - We have rolled out the mitigation and are seeing recovery.

Jul 7, 15:50 UTC
Update - Customers will continue to see 500 errors for approximately 8% of Actions runner REST APIs and 13% of Codespaces REST APIs. Retries may be successful. We have identified a likely cause and are preparing a mitigation.

Jul 7, 15:06 UTC
Update - Customers accessing Actions runners and Codespaces REST APIs continue to see 500 errors a percentage of the time. Retries may be successful.

We continue to investigate the source of these errors.

Jul 7, 14:47 UTC
Update - Customers accessing the Actions and Codespaces REST APIs may see 500 class errors a percentage of the time.  Retries may be successful.

Actions runs and codespaces  in progress are continuing successfully.

We are continuing to investigate the source of these errors.

Jul 7, 14:32 UTC
Update - Codespaces is experiencing degraded availability. We are continuing to investigate.

Jul 7, 14:14 UTC
Update - Customers accessing the Actions and Codespaces REST APIs may see 500 class errors a small percentage of the time.

Actions runs in progress are continuing successfully.

Jul 7, 14:14 UTC
Investigating - We are investigating reports of degraded performance for Actions and Codespaces

View →
InngestJul 7, 2026 at 06:45 UTC

Function execution down for a small subset of customers

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Event API (Operational)
  • Inngest Dashboard (Operational)
  • Function execution (Operational)
  • API (REST and GraphQL) (Operational)
  • Observability (Operational)
View →
ResendJul 3, 2026 at 11:21 UTC

Email sending delay

Status: Resolved

We have resolved the underlying issue and services have been recovered.

Affected components
  • Single Email (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
View →
GitHubresolvedJul 2, 2026 at 16:54 UTC

Incident with Pages

Jul 2, 18:25 UTC
Resolved - On July 2nd, 2026, between approximately 15:00 and 18:30 UTC, the GitHub Pages service experienced degraded deployment performance due to a surge in demand that exceeded available processing capacity. During this period, users publishing to GitHub Pages may have seen their deployments queued or taking substantially longer than usual to go live. No other GitHub services were impacted.

We mitigated the incident by scaling up Pages deployment workers and provisioning additional storage capacity to clear the backlog.

GitHub is reviewing capacity planning and autoscaling measures to reduce the likelihood of similar delays in the future.

Jul 2, 17:57 UTC
Update - Pages deployment latency is recovering. The team continues working toward full mitigation and a return to nominal state.

Jul 2, 16:56 UTC
Update - We are investigating reports of slow and failing Pages deployments. Access to Pages is unaffected.

Jul 2, 16:54 UTC
Investigating - We are investigating reports of degraded performance for Pages

View →
VercelresolvedJul 2, 2026 at 16:33 UTC

Degraded performance on CDN, Dashboard, Functions in Washington, Cleveland (iad1, cle1)

Jul 2, 16:46 UTC
Resolved - This incident has been resolved.

Jul 2, 16:39 UTC
Monitoring - A fix has been implemented and we are observing recovery across CDN, Dashboard, & Functions in Washington, Cleveland (iad1, cle1). We are continuing to monitor.

Jul 2, 16:33 UTC
Identified - The issue has been identified and a fix is being implemented.

Jul 2, 16:31 UTC
Investigating - We are investigating an issue causing degraded performance across CDN, Dashboard, & Functions in Washington, Cleveland (iad1, cle1). We'll provide additional updates as they become available.

View →
VercelresolvedJul 1, 2026 at 16:57 UTC

Increase in Workflow runs stuck in pending, failed

Jul 1, 18:30 UTC
Resolved - The issue affecting workflow runs getting stuck in a pending or failed state has been resolved. New workflow runs after the impact period are unaffected.

Affected workflow runs have been recovered where possible. Runs that could be resumed safely were resumed and runs that could not be resumed safely were transitioned to a cancelled state.

Jul 1, 17:49 UTC
Update - We are continuing to work on recovering affected workflow runs that may be stuck in a pending state or have reached a failed state.

New workflow runs are not impacted.

Jul 1, 16:57 UTC
Identified - We've identified an issue where some workflow runs created between June 30 at 22:24 UTC and July 1 at 16:08 UTC may be stuck in a pending state or have reached a failed state.

New workflow runs are not impacted. We are working on a fix for runs stuck in a pending state and will provide additional updates as they become available.

View →
ResendJul 1, 2026 at 16:00 UTC

Elevated error rates across our APIs and SMTP services

Status: Resolved

Our team has mitigated the issue and this incident is fully resolved. All services are operational.

Affected components
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Website (Operational)
  • Automations (Operational)
  • Single Email (Operational)
  • Dashboard (Operational)
  • Webhooks (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
View →
GitHubresolvedJul 1, 2026 at 10:51 UTC

Delays in copilot budget limits resets for some users

Jul 1, 13:26 UTC
Resolved - On July 1, 2026, between approximately 00:00 UTC and 13:04 UTC, some GitHub Copilot customers whose budget was exhausted before the monthly reset remained incorrectly blocked from paid Copilot usage after the new billing month began, even though their budgets had reset. Some budget changes also took longer than usual to apply. Only customers with an exhausted budget were affected, which limited the impact.

This was caused by a caching issue at the monthly reset: for some users, a pre-reset "budget exhausted" status was re-saved and served even though their budget had reset, so they stayed blocked. We had built a safeguard ahead of the reset to prevent this, but it did not take effect because an internal configuration service did not load its settings correctly. We resolved the incident by deploying a change that discards the outdated status and recomputes access from current budget data independently of that configuration, and by working through the backlog of budget updates.

To prevent recurrence, we are ensuring pre-reset status cannot survive the monthly budget reset, adding alerting for this failure mode, and increasing capacity to absorb the monthly surge of budget updates.

Jul 1, 13:05 UTC
Update - The fix has been deployed globally and we are monitoring the results

Jul 1, 13:04 UTC
Monitoring - The degradation affecting Copilot has been mitigated. We are monitoring to ensure stability.

Jul 1, 12:42 UTC
Update - A fix for the delayed resets is currently being deployed.

Jul 1, 11:44 UTC
Update - We have identified the likely reason for the delays and are working on a solution.

Jul 1, 10:51 UTC
Investigating - We are investigating reports of degraded performance for Copilot

View →
GitHubresolvedJun 30, 2026 at 15:38 UTC

Disruption with some GitHub services - Signup Flow

Jun 30, 15:49 UTC
Resolved - Between 15:19 UTC and 15:49 UTC on June 30, 2026, users were unable to complete the signup flow for GitHub.com/signup. Approximately 62% of new user signups failed for about 30 minutes during this window.

This was caused by a configuration change to the signup flow that unintentionally blocked users from completing signup.

We mitigated the incident by reverting the change, which restored successful signups. To reduce the likelihood and impact of similar issues, we are adopting staged, incremental rollouts for changes on the signup path, improving our ability to test these changes before they reach production, and adding checks to verify signup health before and during any change that affects this flow.

Jun 30, 15:38 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 28, 2026 at 17:50 UTC

Disruption with some GitHub services

Jun 28, 20:55 UTC
Resolved - From June 26, 2026 at 23:40 UTC through June 28, 2026 at 20:55 UTC, Copilot Cloud Agent was degraded. The agent could fail when reporting progress, replying to pull request comments, or opening pull requests. For affected built-in tool calls, the average error rate was approximately 8%, with hourly error rates peaking around 26%.

This was due to a regression introduced during a Copilot Cloud Agent runtime deployment that caused several built-in agent tools to become unavailable. In many cases, the affected tool calls failed silently so agent jobs appeared to succeed. This monitoring gap meant it took longer than expected to identify the failure. We mitigated the incident by reverting the runtime deployment to the previously stable version.

We've added monitoring and alerting for this class of tool-availability error to reduce time-to-detection. We're also adding regression tests for these built-in agent tools, and improving the shipping safety for future runtime rollouts to avoid similar issues.

Jun 28, 20:02 UTC
Update - Copilot cloud agent had been experiencing intermittent problems with opening pull requests, pushing changes and replying to comments. A fix has been deployed and we are validating that fix.

Jun 28, 18:31 UTC
Update - Copilot cloud agent is experiencing intermittent problems with opening pull requests, pushing changes and replying to comments. We have validated a fix and are deploying that fix now.

Jun 28, 17:59 UTC
Update - Copilot cloud agent is experiencing intermittent problems with opening pull requests, pushing changes and replying to comments. We have identified the issue and are validating a fix.

Jun 28, 17:50 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 27, 2026 at 14:04 UTC

Disruption with some GitHub services

Jun 27, 20:33 UTC
Resolved - This incident was used to notify for a maintenance event. There is no specific root cause analysis. Work progressed as planned without any issues to report.

Jun 27, 20:19 UTC
Update - Maintenance has completed and we've begun normalizing traffic in EU. We will monitor for a while longer before confirming resolution.

Jun 27, 18:22 UTC
Update - This is progressing as expected, but we're going to extend the maintenance window by a couple of hours. Our new expected completion time is 21:00 UTC.

Jun 27, 14:04 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 27, 14:02 UTC
Update - We are conducting routine maintenance on our network infrastructure in the EU. This will not impact production traffic, but may result in slightly increased latency for the remainder of our work. We expect this to last until 19:00 UTC.

Jun 27, 14:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedJun 26, 2026 at 13:08 UTC

Intermittent Login Failures on Vercel Dashboard

Jun 26, 15:14 UTC
Resolved - This incident has been resolved.

Jun 26, 15:03 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 26, 14:29 UTC
Update - Login is recovering for most users. We're monitoring to ensure stability.

Jun 26, 14:11 UTC
Identified - The issue has been identified and a fix is being implemented.

Jun 26, 14:08 UTC
Update - We are continuing to investigate this issue and will provide an update as soon as we have more information.

Jun 26, 13:23 UTC
Update - We are continuing to investigate this issue.

Jun 26, 13:08 UTC
Investigating - We are currently investigating the issue. Some users are experiencing errors logging in. The issue started at 11:30 AM UTC.

View →
GitHubresolvedJun 25, 2026 at 17:50 UTC

Degradation with Webhooks, Pull Requests and Actions

Jun 25, 18:27 UTC
Resolved - On June 25, 2026, between 17:33 UTC and 17:55 UTC, our background job service experienced degradation which increased delays to pull requests, repository pushes, Actions workflows, and Webhooks, with delays peaking at 7m. The issue was caused by underlying hypervisor issues and an incoming traffic spike, causing service timeouts which led to a connection storm and continual rebalances.

The issue was mitigated by replacing the problem node at 17:49, after which all services saw recovery by 18:07.

Jun 25, 18:27 UTC
Update - We identified an issue that caused degradation across multiple services including Webhooks, Pull Requests, Actions, and Issues. Customers may have experienced delays or failures with these services. We have applied mitigations and affected services have recovered.

Jun 25, 18:07 UTC
Monitoring - The degradation affecting Actions, Issues, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

Jun 25, 17:58 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Jun 25, 17:53 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

Jun 25, 17:50 UTC
Investigating - We are investigating reports of degraded performance for Actions, Pull Requests and Webhooks

View →
VercelresolvedJun 25, 2026 at 07:45 UTC

Increase in Workflow runs stuck as pending

Jun 25, 13:43 UTC
Resolved - This incident has been resolved.

Jun 25, 13:30 UTC
Monitoring - The fix has been fully rolled out and we are monitoring the results.

Jun 25, 12:15 UTC
Update - We are currently rolling a fix for the issue. We are monitoring the results.

Jun 25, 10:42 UTC
Update - We are still working on a fix for this issue. While we fix the issue, users experiencing pending Workflow runs can trigger an instant rollback to a working deployment created before Jun 24 21:00 UTC to resume operation on new runs.

Jun 25, 09:25 UTC
Update - The issue has been identified. We are continuing to work on a fix for this issue.

Jun 25, 08:27 UTC
Identified - The issue has been identified and a fix is being implemented.

Jun 25, 07:45 UTC
Investigating - We are currently investigating an issue where some recently deployed Workflow projects are stuck as pending and not delivering queued messages as expected.

View →
VercelresolvedJun 23, 2026 at 23:24 UTC

Degraded Domain Search in Vercel Dashboard

Jun 23, 23:34 UTC
Resolved - This incident has been resolved.

Jun 23, 23:28 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 23, 23:26 UTC
Identified - The issue has been identified and a fix is being implemented.

Jun 23, 23:24 UTC
Investigating - We're investigating an issue preventing customers from searching for and purchasing domains in the Vercel Dashboard.

View →
GitHubresolvedJun 23, 2026 at 23:04 UTC

We are seeing elevated errors with Next Edit Suggestions and Completions

Jun 23, 23:29 UTC
Resolved - On June 23, 2026, between 22:45 and 23:29 UTC, GitHub Copilot Completions and Next Edit Suggestions were degraded for users in all regions. During this window, affected users may have seen failed or missing code completions and Next Edit Suggestions. On average about 25% of Completions and Next Edit Suggestions requests failed during the impact window, peaking at roughly 27%. The cause was a configuration change that prevented the Copilot service from obtaining the authentication tokens it needs to reach its model backends; this both failed requests directly and caused the service to temporarily remove backends from rotation. GitHub engineers detected the elevated error rate within minutes, declared an incident, and mitigated the issue at 23:22 UTC by redeploying the service with a known-good configuration, which restored normal operation. As a follow-up, the team disabled the affected authentication path to prevent a future deployment from re-introducing the problem, and is making the change rollout safer. We apologize for the disruption and are taking steps to reduce the likelihood of similar incidents.

Jun 23, 23:27 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 23, 23:26 UTC
Update - We have applied a rollback and are seeing recovery.

Jun 23, 23:04 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendJun 22, 2026 at 22:07 UTC

Degraded perfomance in Metrics

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Dashboard (Operational)
View →
ResendJun 22, 2026 at 20:57 UTC

Delayed webhooks

Status: Resolved

We have resolved the underlying issue, and all the affected webhooks have been sent.

Affected components
  • Webhooks (Operational)
  • Email Events (Operational)
View →
VercelresolvedJun 19, 2026 at 17:54 UTC

Resolved: Elevated ERR_STREAM_PREMATURE_CLOSE errors in Vercel Functions

Jun 19, 17:54 UTC
Resolved - Between Jun 19 09:09 and 16:16 UTC, a subset of Vercel Functions using node-fetch@2 may have experienced intermittent invocation errors that surfaced as ERR_STREAM_PREMATURE_CLOSE.

As part of Node.js June 2026 security releases, we began rolling out new Node.js versions. Those versions contain an upstream regression that breaks response streaming for node-fetch@2, which caused the errors https://github.com/nodejs/node/issues/63989

We have resolved the issue by reverting to the previous Node.js version. No action is required — affected functions are now operating normally. We will re-land the Node.js upgrade once the upstream issue is fixed.

View →
InngestJun 19, 2026 at 13:07 UTC

Support Portal Unavailable

Status: Resolved

The Support Portal is now available.

Affected components
  • Inngest Dashboard (Operational)
View →
GitHubresolvedJun 17, 2026 at 19:00 UTC

Incident With Webhooks

Jun 17, 19:00 UTC
Resolved - On June 17, 2026, between 11:35 UTC and 19:20 UTC, the Webhooks service was degraded and delivered webhook payloads with missing installation information. On average, 11.3% of webhook deliveries were impacted. Customers relying on the installation field for authentication or routing were unable to process affected webhooks. A smaller subset of deliveries for the security_advisory event (0.04%) were delivered successfully but were not recorded for redelivery. This was due to a defect in a new delivery code path that failed to include installation data in webhook payloads.

We mitigated the incident by disabling the feature flag controlling the new code path.

We are working to improve our automated validation of webhook payloads, and introduce automated alerting for webhook payload regressions to reduce our time to detection and mitigation of issues like this one in the future.

The following events were affected: branch_protection_configuration, code_scanning_alert, commit_comment, custom_property, custom_property_values, dependabot_alert, deploy_key, deployment_protection_rule, deployment_review, dismissal_request_code_scanning, dismissal_request_secret_scanning, installation_target, member, membership, merge_queue_entry, org_block, organization, projects_v2, projects_v2_item, pull_request_review_thread, repository_ruleset, secret_scanning_alert, secret_scanning_alert_location, secret_scanning_scan, security_and_analysis, star, sub_issues, team, team_add, workflow_job.

View →
GitHubresolvedJun 17, 2026 at 17:57 UTC

Disruption with Copilot next edit suggestions

Jun 17, 19:28 UTC
Resolved - On June 17, 2026, between 16:57 UTC and 19:14 UTC, Copilot code completions were degraded and users were unable to receive Next Edit Suggestions. Standard ghost text code completions were not affected. This was due to a configuration change that caused the service's routing layer to incorrectly discard all Next Edit Suggestion model endpoints as invalid.

We mitigated the incident by deploying a corrected configuration change at 18:55 UTC, with full recovery observed at 19:14 UTC.

We are working to improve the resilience of our routing layer to limit impact due to a subset of invalid configurations, and to improve our alerting to detect sudden traffic changes that are not captured by standard error rate monitors.

Jun 17, 19:14 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 17, 19:09 UTC
Update - We have applied mitigation and are seeing recovery

Jun 17, 18:40 UTC
Update - We have identified a likely cause and are applying a mitigation

Jun 17, 17:57 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendJun 17, 2026 at 16:55 UTC

Automations not starting

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • Automations (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Single Email (Operational)
View →
ResendJun 17, 2026 at 16:29 UTC

Email Events partially unavailable

Status: Identified

We have identified the issue and we are working on a fix now.

Affected components
  • Email Events (Partial outage)
View →
GitHubresolvedJun 17, 2026 at 03:52 UTC

Incident with Copilot Availability

Jun 17, 04:44 UTC
Resolved - On June 17, 2026, between approximately 03:35 UTC and 04:44 UTC, GitHub Copilot was degraded and most of its frontier chat models were temporarily unavailable across all regions. During this window, affected models either disappeared from the model picker in the web, editor, and CLI experiences, or returned a "model not available" error when selected. Customers could continue using GitHub Copilot by selecting one of the models that remained available. The incident occurred during off-peak hours, which limited the number of customers affected.

This was due to a configuration change that our production system deemed invalid. We mitigated the incident by reverting the configuration change, after which the affected models returned automatically as the service reloaded the previous configuration.

We are working to roll out configuration changes gradually with stronger validations, alerts on sudden drops in the number of available models, and automatically rolls back configuration changes that produce these alerts.

Jun 17, 04:44 UTC
Monitoring - Copilot is operating normally.

Jun 17, 04:26 UTC
Update - We've applied a mitigation to unblock Copilot functionality. Users may start to see signs of recovery. Relaunching your client should accelerate signs of recovery. We will continue to monitor the situation.

Jun 17, 03:52 UTC
Update - We are experiencing degraded availability for chat & agent models in Copilot. Multiple models are impacted and customers may experience requests failing. We are investigating and will provide an update as soon as possible.

Jun 17, 03:50 UTC
Investigating - We are investigating reports of degraded availability for Copilot

View →
GitHubresolvedJun 16, 2026 at 17:47 UTC

Disruption with some GitHub services

Jun 16, 18:15 UTC
Resolved - On June 16, 2026, between 17:20 UTC and 18:15 UTC, the Opus 4.8 model experienced degraded availability in GitHub Copilot. During this window, some requests to Opus 4.8 failed or errored. Other Copilot models were not affected and remained available as alternatives. This was caused by an issue with an upstream model provider. The upstream provider resolved the issue, and we monitored Opus 4.8 until success rates returned to normal. The incident is fully resolved.

Jun 16, 18:14 UTC
Update - The issues with our upstream model provider have been resolved, and Opus 4.8 is once again available in Copilot products and IDE surfaces.

We will continue monitoring to ensure stability, but mitigation is complete.

Jun 16, 18:00 UTC
Monitoring - The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

Jun 16, 17:47 UTC
Update - We are experiencing degraded availability for the Opus 4.8 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Jun 16, 17:45 UTC
Investigating - We are investigating reports of degraded availability for Copilot AI Model Providers

View →
VercelresolvedJun 16, 2026 at 13:39 UTC

Partial disruption of Observability, Analytics, Firewall, and Usage in dashboard

Jun 16, 15:39 UTC
Resolved - The incident has been resolved.

Jun 16, 15:26 UTC
Monitoring - The team has deployed a fix for timeouts and failed requests related to Observability, Web Analytics, Speed Analytics, Firewall, and Usage in dashboard. We are continuing to monitor.

Jun 16, 15:16 UTC
Identified - The team has identified the source of the issue and is preparing a fix for timeouts and failed requests related to Observability, Web Analytics, Speed Analytics, Firewall, and Usage in dashboard.

Jun 16, 14:45 UTC
Update - We are continuing to investigate partial disruption to Observability, Web Analytics, Speed Analytics, Firewall, and Usage in dashboard. We are also investigating delayed usage alerts.

Jun 16, 13:39 UTC
Investigating - We are investigating a partial disruption to Observability, Web Analytics, Speed Analytics and Firewall features.

View →
ResendJun 15, 2026 at 21:14 UTC

Issues downloading attachments

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • General API (Operational)
View →
GitHubresolvedJun 15, 2026 at 18:32 UTC

Multiple services have elevated errors and endpoint failures when checking feature flags

Jun 15, 19:10 UTC
Resolved - Between 17:38 UTC and 18:22 UTC on June 15, 2026, approximately 83% of requests to the analytics endpoint serving the /chronicle feature failed.  The cause was an internal feature-flag service that encountered a transient error and failed to recover, causing feature flag checks to fail. The analytics endpoint was gated behind one of these flags, resulting in requests being rejected. We restored service health by removing the feature flag gating the analytics endpoint and deploying that change. To avoid recurrence of similar incidents, we have changed the feature-flag client so that errors that are not known to be permanent are retried, and we are improving alerting and startup behavior so this class of failure is detected and recovered from faster.

Jun 15, 19:10 UTC
Update - We identified an issue with feature flag checks that caused elevated errors and endpoint failures across multiple GitHub Copilot services. Customers may have experienced failed requests or degraded functionality. A fix has been deployed and error rates are steadily decreasing. All affected services are now mitigated and we are monitoring recovery.

Jun 15, 19:02 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 15, 18:32 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 15, 2026 at 15:37 UTC

Increased latency with webhooks

Jun 15, 17:37 UTC
Resolved - On June 15, 2026, between 15:27 UTC and 16:23 UTC, GitHub webhook deliveries were delayed. During this window, webhook events were delivered later than normal, with average end-to-end delivery latency peaking at approximately 8.8 minutes. No webhook deliveries were lost — delayed events were queued and delivered once processing recovered.

This was caused by a temporary throughput constraint in an internal event-processing system that moves webhook events through GitHub's delivery pipeline. The rate at which events were processed for delivery dropped below the incoming volume, creating a backlog. We restarted the affected pipeline service, after which throughput recovered and the backlog fully drained by approximately 16:29 UTC. Webhook delivery latency returned to normal, the incident was mitigated at 16:39 UTC, and fully resolved at 17:37 UTC.

To reduce the likelihood and impact of similar incidents, we are working on improving the accuracy of the utilization metrics used to scale our delivery worker pools, reviewing connection and capacity headroom in the delivery pipeline.

Jun 15, 16:39 UTC
Monitoring - The degradation affecting Webhooks has been mitigated. We are monitoring to ensure stability.

Jun 15, 16:39 UTC
Update - Webhooks delivery latency has returned to normal levels. The backlog that built up during the incident has been cleared at approximately 4:29 UTC. We consider this incident resolved.

Jun 15, 15:37 UTC
Investigating - We are investigating reports of degraded performance for Webhooks

View →
VercelresolvedJun 14, 2026 at 21:03 UTC

Elevated latency and errors in FRA1 Vercel Region (Frankfurt, Germany)

Jun 15, 01:24 UTC
Resolved - This incident has been resolved.

Jun 15, 00:25 UTC
Monitoring - All traffic has been restored to the FRA1 Vercel Region.

Jun 14, 23:15 UTC
Update - The issue has been fixed. We are starting to re-route traffic from the CDG1 Vercel Region back to the FRA1 Vercel Region. We will provide another status update after this has finished.

Jun 14, 21:26 UTC
Update - The impact is currently mitigated. We are continuing to re-route traffic from the FRA1 Vercel Region to the CDG1 Vercel Region.

Jun 14, 21:07 UTC
Identified - We are currently re-routing traffic from the FRA1 Vercel Region to the CDG1 Vercel Region.

Jun 14, 21:03 UTC
Investigating - We are currently re-routing traffic from the FRA1 Vercel Region to the CDG1 Vercel Region. We are currently investigating this issue.

View →
GitHubresolvedJun 11, 2026 at 19:42 UTC

Incident with Webhooks

Jun 11, 22:19 UTC
Resolved - On June 11, 2026, between 19:28 UTC and 21:06 UTC, GitHub webhook deliveries were delayed. Average delivery latency peaked at approximately 3.4 minutes, with some deliveries delayed by as much as 62 minutes at the 99th percentile. No events were lost — delayed events were queued and delivered once processing caught up.

This was due to a change in how webhook traffic was distributed across regions: to relieve load on one region, a portion of processing was shifted to another, where higher latency prevented our delivery workers from keeping pace with incoming volume, creating a backlog. We mitigated the incident by rebalancing webhook traffic distribution; as load returned to normal levels, processing caught up and the delivery backlog fully drained.

We are working on improving the accuracy of the utilization metrics used to scale our delivery worker pools, and reassess how we distribute webhook traffic across regions, to reduce our time to detection and mitigation of issues like this one in the future.

Jun 11, 20:00 UTC
Monitoring - The degradation affecting Webhooks has been mitigated. We are monitoring to ensure stability.

Jun 11, 19:58 UTC
Update - We have applied a mitigation and are monitoring for recovery.

Jun 11, 19:42 UTC
Update - We are currently experiencing delays in Web Hook delivery and are actively investigating the root cause.

Jun 11, 19:42 UTC
Investigating - We are investigating reports of degraded performance for Webhooks

View →
GitHubresolvedJun 10, 2026 at 15:24 UTC

Authentication issues related to API requests

Jun 10, 16:39 UTC
Resolved - Between 15:05 UTC and 16:25 UTC, GitHub API services experienced degraded availability due to sporadic authentication failures affecting approximately 9% of requests. Customers experienced intermittent "logged out" behavior as erroneous 401 responses triggered repeated authentication flows in app integrations. Affected requests also experienced approximately 800ms of additional latency.

A memcached proxy service rollout to our internal API infrastructure caused our authentication service to pick up an incorrect memcached host configuration, leading to intermittent authentication lookup failures. We mitigated the incident by deploying a configuration change to memcached to use the correct host.

To prevent similar issues in the future, we plan to migrate our authentication system to the new memcached infrastructure to improve resilience and strengthen overall reliability posture.

Jun 10, 16:37 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 10, 16:36 UTC
Update - The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

Jun 10, 16:21 UTC
Update - We continue to investigate issues related to sporadic authentication failures, impacting approximately 15% of API traffic. Erroneous 401 responses are causing app integrations to trigger authentication flows. We have identified a problematic component in our infrastructure and are working to mitigate.

Jun 10, 15:46 UTC
Update - The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

Jun 10, 15:46 UTC
Update - We continue to investigate issues related to sporadic authentication failures, impacting approximately 15% of API traffic. Further updates will be provided as we work to mitigate.

Jun 10, 15:27 UTC
Update - We are investigating issues related to sporadic authentication failures impacting approximately 15% of API traffic. We will continue to investigate and provide updates.

Jun 10, 15:27 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Jun 10, 15:23 UTC
Update - API Requests is experiencing degraded availability. We are continuing to investigate.

Jun 10, 15:20 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedJun 9, 2026 at 14:09 UTC

Elevated Build Errors for Secure Compute/Static IPs Projects

Jun 9, 14:46 UTC
Resolved - This incident has been resolved.

Jun 9, 14:39 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 9, 14:33 UTC
Identified - The issue has been identified and a fix is being implemented.

Jun 9, 14:09 UTC
Investigating - We are currently investigating this issue.

View →
VercelresolvedJun 8, 2026 at 18:57 UTC

Elevated Functions Invocation Errors in DUB1 (Dublin, Ireland) Region

Jun 8, 20:00 UTC
Resolved - This incident has been resolved.

Jun 8, 19:36 UTC
Update - We are continuing to monitor for any further issues.

Jun 8, 19:36 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 8, 19:04 UTC
Update - We're continuing to work on a fix. In the meantime, deployments with multiple function regions or failover regions are being rerouted to the nearest healthy region.

If your project uses only the dub1 function region, you can switch to the nearest region, lhr1, and redeploy to mitigate the issue.

Jun 8, 18:57 UTC
Identified - The issue has been identified and a fix is being implemented.

A small number of requests may have seen elevated error rates in function invocations during this period.

View →
GitHubresolvedJun 8, 2026 at 15:00 UTC

Degraded availability for GitHub.com, GraphQL API, and Webhooks UI/API

Jun 8, 15:00 UTC
Resolved - On June 8, 2026, between 14:49 and 14:54 UTC, a subset of requests to GitHub.com, the REST API, GraphQL API, and Webhooks UI/API experienced elevated error rates due to a transient infrastructure capacity issue that self-resolved within approximately 5 minutes.

Users experienced HTTP 500 errors and timeouts when accessing GitHub.com, the REST API, GraphQL API, and Webhooks UI/API for approximately 5 minutes, with the REST API taking up to 12 minutes to fully recover.

View →
GitHubresolvedJun 8, 2026 at 09:08 UTC

Disruption with Claude Opus 4.7

Jun 8, 10:03 UTC
Resolved - On June 8, 2026, between 08:40 UTC and 09:30 UTC, the Claude Opus 4.7 model experienced degraded availability with error rates peaking at 8.4% and averaging 1.9%. This was due to an upstream provider issue that caused temporary unavailability and rate limiting on secondary failover systems. Users selecting Auto or alternative models were unaffected. We are improving provider failover mechanisms and monitoring to prevent similar issues.

Jun 8, 09:49 UTC
Update - We are experiencing degraded availability for the Claude Opus 4.7 model in Copilot Chat, VS Code and other Copilot products. This is due to an issue with an upstream model provider. We are working with them to resolve the issue, and starting to see recovery.

Jun 8, 09:08 UTC
Update - We are experiencing degraded availability for the Claude Opus 4.7 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

Jun 8, 09:05 UTC
Investigating - We are investigating reports of degraded performance for Copilot AI Model Providers

View →
GitHubresolvedJun 8, 2026 at 07:14 UTC

Pull Requests and Issues unavailable for signed-out users

Jun 8, 08:36 UTC
Resolved - On June 8, 2026, between approximately 06:30 UTC and 08:36 UTC, signed-out users experienced sustained elevated HTTP 504 errors when accessing Pull Requests, Issues, releases, patch diffs, and other related GitHub.com pages. During the incident, approximately 17% of unauthenticated requests to the affected GitHub.com endpoints returned gateway timeout errors, peaking at roughly 34% of requests at around 06:50 UTC. Some GitHub Actions workflows were also affected when they depended on release downloads or related GitHub.com endpoints. The impact lasted approximately two hours and was isolated to unauthenticated traffic; signed-in users were not affected.

The issue was caused by a significant increase in abusive traffic to specific GitHub.com endpoints. This degraded our ability to respond to unauthenticated requests, causing requests to queue beyond timeout thresholds and return gateway timeout errors.

We mitigated the incident by identifying the anomalous traffic pattern and applying targeted blocks at the load balancer and application layers. Error rates returned to normal and affected services were fully restored by 08:36 UTC.

To reduce the likelihood and impact of similar incidents in the future, we are improving automated detection and blocking for these traffic patterns, improving our emergency traffic-blocking deployment path, and evaluating routing changes for endpoints used by both signed-out users and automated workflows.

Jun 8, 08:35 UTC
Monitoring - The degradation affecting Actions, Issues and Pull Requests has been mitigated. We are monitoring to ensure stability.

Jun 8, 08:27 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Jun 8, 08:13 UTC
Update - Following investigation, we are seeing that impact is limited to unauthenticated users when accessing Pull Requests, Issues, or Actions. Our team continues to work towards mitigation with more updates to follow as we have them.

Jun 8, 07:32 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Jun 8, 07:31 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Jun 8, 07:14 UTC
Update - Issues is experiencing degraded availability. We are continuing to investigate.

Jun 8, 07:11 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 6, 2026 at 16:53 UTC

Disruption with some GitHub services in the EU region

Jun 6, 17:07 UTC
Resolved - On June 6, 2026 between 16:18 UTC and 17:01 UTC, users experienced elevated error rates when performing Git operations (cloning, fetching, downloading archives) and accessing package registries. The issue affected users whose traffic was routed through our European infrastructure.

During this time, on average 0.95% of Codeload requests and 9.2% of Package Registry requests failed with server errors. At peak, the Codeload error rate reached 1.76% and Package Registry errors reached 27%.

The root cause was a planned network circuit migration that disrupted connectivity at one of our points of presence. Our process for shifting traffic away from the site did not operate as expected, resulting in a small amount of production traffic to continue being serviced at the effected site during the maintenance window. The issue was mitigated by rolling back the network change, restoring normal connectivity. Services fully recovered by 17:01 UTC.

To reduce the likelihood of similar incidents in the future, we are reviewing our site drain process to make it more verbose and add visibility so any unexpected behavior is caught earlier.

Jun 6, 16:56 UTC
Update - Packages is experiencing degraded performance. We are continuing to investigate.

Jun 6, 16:53 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 6, 2026 at 15:31 UTC

EU Network Maintenance

Jun 6, 18:49 UTC
Resolved - This incident was used to notify for a maintenance event. There is no specific root cause analysis. Maintenance did run longer than expected (we were complete at 18:48 UTC) but the work proceeded as planned.

Jun 6, 18:48 UTC
Update - All work has been completed and we are hands off.

Jun 6, 15:36 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 6, 15:31 UTC
Update - We are conducting routine maintenance on our network infrastructure in the EU. This will not impact production traffic, but may result in slightly increased latency for the remainder of our work. We expect this to last until 17:00 UTC.

Jun 6, 15:31 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 5, 2026 at 17:20 UTC

Auth issue resulting in API impacts, including some Slack and Teams channel subscriptions

Jun 5, 22:21 UTC
Resolved - On June 5, 2026, between 15:35 UTC and 16:45 UTC, 0.11% of authenticated REST API requests incorrectly returned “not found” responses. Impact was concentrated among - and significantly higher for - users authenticating with user-to-server tokens to access organization-owned repositories.

Some users of our GitHub for Slack and GitHub for Microsoft Teams integrations saw their channel subscriptions removed as those systems interpreted the transient "not found" response as durable loss of access. Roughly 12% of organizations with active channel subscriptions were impacted, with ~2% of all channel subscriptions being removed.

These issues were triggered by a change to an internal authorization component that did not correctly resolve access for user-to-server tokens against organization-owned repositories. We mitigated the incident by disabling the accompanying feature flag at 16:45 UTC, after which API responses returned to normal. We then restored all impacted Slack and Microsoft Teams channel subscriptions, with restoration completed at 22:21 UTC.

We are working to add retry and grace-period logic in the chat integrations so transient errors no longer trigger subscription deletions. In parallel, we are improving observability and gating of authorization changes so downstream impact is detected during scoped, gradual rollouts.

Jun 5, 22:21 UTC
Update - Affected Slack and Teams subscriptions have been restored. Please contact support if you encounter any additional issues.

Jun 5, 20:34 UTC
Update - Additional detail on the scope of impact during the 14:49 UTC to 16:45 UTC window: a small but elevated percentage of authenticated requests to GitHub.com received incorrect authorization failures. We saw a 1 to 2% increase in 4xx responses for a small number of endpoints (/repos/{owner}/{repo}/pulls/{pull_number}, /repos/{owner}/{repo}, /repos/{owner}/{repo}/contents/{path}). The vast majority of requests completed normally; customers who saw errors during the window can retry now and should see them succeed.

Jun 5, 18:43 UTC
Update - We are still exploring options to restore the deleted subscriptions, and we will provide another update soon. In the meantime, customers can manually re-subscribe their Slack and Teams channels to repositories.

Jun 5, 18:05 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 5, 18:04 UTC
Update - During 14:49 UTC to 16:45 UTC, customers may have experienced authorization failures for legitimate requests. This was caused by a recently enabled feature flag, which has now been turned off as a mitigation. Customers should now see normal authorization behavior. This is also the cause of the chat integration issue, and we are exploring options to restore it. In the meantime, customers can manually re-subscribe their repo.

Jun 5, 17:25 UTC
Update - Customers may see unexpected repo unsubscription events in their Slack or Teams channels.

Jun 5, 17:20 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 4, 2026 at 20:20 UTC

Live updates degraded

Jun 4, 20:32 UTC
Resolved - Everything is operating normally.

Jun 4, 20:20 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 4, 2026 at 18:03 UTC

Copilot Code Review Failing

Jun 4, 19:59 UTC
Resolved - On June 4, 2026, from 17:30 UTC to 18:55 UTC, Copilot Code Review experienced elevated failures for review requests on GitHub.com. Affected users saw “Copilot ran into an error” on pull requests when requesting a code review.

During the incident window, an average of 81.6% of Copilot Code Review requests failed, with a peak failure rate of 93.9%. Approximately 36,800 code review requests failed. GitHub Enterprise Cloud with data residency was not impacted.

The issue was caused by a newly released dependency used by the Copilot Code Review processing workflow. The release introduced an incompatibility with the runtime environment. Because the workflow automatically consumed the latest release, the incompatible version was picked up without sufficient compatibility validation and caused review processing to fail.

We mitigated the incident by removing the problematic dependency version and redeploying the affected processing service. New code reviews began recovering at 18:44 UTC, and the failure rate returned to baseline by 18:55 UTC. Remaining timed-out work drained by 19:59 UTC.

To reduce the risk of recurrence, we are pinning the dependency version instead of automatically consuming the latest release, adding compatibility checks for future releases, improving fast-failure behavior when the review processor cannot start, adding shorter timeout controls for review workflows, and improving monitoring for review completion failures.

Jun 4, 19:59 UTC
Update - This issue is now fully resolved and Copilot Code Review is working as expected.

Jun 4, 19:41 UTC
Update - The mitigation for Copilot Code Review is now fully deployed, and new reviews are working as expected.  We are continuing to monitor for full resolution.

 Customers may need to re-request Copilot Code Review. Copilot Code Review Actions runs running for longer than 20 minutes may be safely cancelled.

Jun 4, 19:07 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 4, 19:07 UTC
Update - The mitigation for Copilot Code Review is now fully deployed, and new reviews are working as expected.

Customers may need to re-request Copilot Code Review.  Copilot Code Review Actions runs running for longer than 20 minutes may be safely cancelled.

Jun 4, 18:52 UTC
Update - The mitigation for Copilot Code Review is rolling out and we are seeing early signs of recovery.

Jun 4, 18:22 UTC
Update - We have identified that Copilot Code Review users may see "Copilot ran into an error" on Pull Requests that requested Copilot Code Review.

A mitigation is in progress, we expect mitigation in approximately 30m.

GitHub Enterprise Cloud with Data Residency is not impacted.

Jun 4, 18:03 UTC
Update - We have identified that Copilot Code Review.  Users may see "Copilot ran into an error" on Pull Requests that requested Copilot Code Review.

A mitigation is in progress.

Jun 4, 18:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedJun 3, 2026 at 19:43 UTC

Disruption with some GitHub services

Jun 4, 04:11 UTC
Resolved - Between June 1, 2026, 23:00 UTC and June 4, 2026 04:11 UTC, customers experienced delays in Dependabot scheduled version updates.

Pull request creation for version updates was delayed, with delays increasing over time and reaching up to two days. Approximately 1.5 million repositories with active Dependabot version update configurations were affected. Dependabot security updates were not affected. The primary cause was changes to an internal platform service that routes requests for Dependabot and other services.

We mitigated the incident by deploying a fix that enables batch enqueuing of update jobs, which significantly increased processing throughput. Once the backlog was drained, Dependabot returned to normal processing times.

To reduce the risk of recurrence, we are working on tuning batch size and concurrency limits for Dependabot update job processing. We are also adding monitoring for job processing lag to enable earlier detection and faster mitigation of similar issues.

Jun 4, 04:11 UTC
Update - Job lag has recovered to within normal operating thresholds. We are declaring this incident closed and will follow up with a summary soon.

Jun 4, 02:37 UTC
Update - Job lag has recovered from a peak of 1.71 days to 9h 9m at 19:29 UTC and continues to decrease. Backlog is draining at a healthy rate with no signs of reversal. New jobs are processing on schedule. Remaining lag will continue to drain over the next few hours as queued work completes; this is expected post-incident catch-up, not active impact. We will continue monitoring and re-engage if lag trend reverses.

Jun 4, 02:35 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 3, 23:38 UTC
Update - We have applied mitigations and are continuing to see improvements in the Dependabot scheduled version updates.

Next update in 12 hours.

Jun 3, 21:10 UTC
Update - We are preparing a mitigation for the delayed Dependabot scheduled version updates.

Next update in 2 hours.

Jun 3, 20:18 UTC
Update - Customers may see delays of up to two days in Dependabot version updates.

Dependabot Security updates are not delayed.

The team is investigating mitigations for the backlog.

Next update in 1 hour.

Jun 3, 19:43 UTC
Update - We're seeing delays in Dependabot scheduled version update runs. Our team is actively working on a fix and will share updates as the situation develops.

Jun 3, 19:42 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendJun 3, 2026 at 16:15 UTC

API quota reporting incorrect numbers

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Batch Emails (Operational)
  • Single Email (Operational)
  • Broadcast Emails (Operational)
View →
GitHubresolvedJun 3, 2026 at 03:13 UTC

Disruption with some GitHub services

Jun 3, 06:46 UTC
Resolved - On June 2, 2026, between 21:54 UTC and June 3, 2026 06:45 UTC, the Spark service was degraded and users were unable to store or retrieve data for their Spark apps in one of our hosting regions. Users could still make changes to their app configuration during this time. The error rate peaked at 25% of affected requests to the service. Impact was limited to users whose requests were served through a single affected region; 43 users experienced errors during this window.

The root cause was a configuration that referenced a service component by a fixed address rather than a dynamic service endpoint. When the component was replaced, requests could no longer reach the fixed address and began to fail. We resolved the incident by updating the configuration to use a our standard service endpoints that are resilient to component replacement. Recovery time was extended because replacing the component required overrides to a temporary deployment safeguard.

We are working to add validation that prevents fixed infrastructure addresses from being used in application configuration outside of test environments and to improve our monitoring to reduce our time to detect.

Jun 3, 06:45 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Jun 3, 06:02 UTC
Update - We are investigating reports of issues with service(s): Spark. We will continue to keep users updated on progress towards mitigation.

Jun 3, 03:45 UTC
Update - We are investigating reports of impacted performance for some GitHub services.

---
Relevant stamps: dotcom

Jun 3, 03:13 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedJun 1, 2026 at 19:21 UTC

Elevated Errors Creating New Deployments

Jun 1, 20:02 UTC
Resolved - This incident has been resolved.

Jun 1, 19:55 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 1, 19:50 UTC
Identified - The issue has been identified and a fix is being implemented.

Jun 1, 19:21 UTC
Investigating - We are investigating reports of some customers experiencing elevated errors creating new deployments. We will provide additional updates as they become available.

View →
GitHubresolvedJun 1, 2026 at 15:17 UTC

Delays with Code Scanning and Billing

Jun 2, 00:17 UTC
Resolved - Starting from 13:00 UTC June 1, 2026, to 00:17 UTC June 2, 2026, multiple services experienced delayed job processing due to increased latency in our background job queue service. The root cause was insufficient queue processing capacity to handle a large week-over-week increase in total job traffic.

Users saw up to 90 minutes of delay in billing usage updates, 30 minutes of delay for webhook notifications to show, and 15 minutes of delay to see email notifications. Mitigation involved scaling up our background job service capacity to handle the spike in job traffic.

We have added queue capacity monitoring to our background job queue service to stay ahead of weekly growth patterns and to reduce time to detect in the future.

Jun 1, 21:59 UTC
Update - We have identified the root cause and applied mitigations to address delays in billing updates and are continuing to see improvement in the processing rate. We will continue to monitor the progress and will provide an update in few hours.

GitHub Enterprise Cloud with Data Residency is not impacted.

Jun 1, 19:27 UTC
Update - We are continuing to investigate delayed billing updates, on GitHub.com.  We have applied additional mitigations are continuing to see signs of improvement, and are continuing to work to improve the processing rate. We will continue to keep users updated on progress towards mitigation.

GitHub Enterprise Cloud with Data Residency is not impacted.

Next update in 2 hours.

Jun 1, 18:48 UTC
Update - We are continuing to investigate delayed billing updates, on GitHub.com.  We have applied additional mitigations and are seeing some more signs of improvement, and are continuing to work to improve the processing rate. We will continue to keep users updated on progress towards mitigation.

GitHub Enterprise Cloud with Data Residency is not impacted.

Jun 1, 17:28 UTC
Update - We are continuing to investigate delayed billing updates, on GitHub.com.   We have applied multiple mitigations and are seeing some signs of improvement, and are continuing to work to improve the processing rate.  We will continue to keep users updated on progress towards mitigation.

GitHub Enterprise Cloud with Data Residency is not impacted.

Jun 1, 16:42 UTC
Update - We are investigating reports of delayed billing updates, on GitHub.com. We are continuing to investigate delays in our job processing architecture. We are attempting to mitigate at the infrastructure level.  Code scanning runs and notifications have recovered.  We will continue to keep users updated on progress towards mitigation.

GitHub Enterprise Cloud with Data Residency is not impacted.

Jun 1, 15:43 UTC
Update - We are investigating reports of delayed code scanning runs, billing updates, email and mobile push notifications. We are investigating delays in our job processing architecture. We will continue to keep users updated on progress towards mitigation.

Jun 1, 15:17 UTC
Update - We are investigating reports of delayed code scanning runs and delayed billing updates.  We will continue to keep users updated on progress towards mitigation.

Jun 1, 15:17 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedMay 28, 2026 at 19:07 UTC

Elevated error rates across multiple services

May 28, 19:07 UTC
Resolved - On May 28, 2026, between 19:07 UTC and 19:16 UTC, multiple GitHub services experienced elevated error rates. This was due to a change that was partially deployed to an authentication service, causing errors for dependent services including the web experience, REST API, Git operations, and GitHub Actions. At peak impact, 10% of GitHub Actions runs failed to queue or encountered errors while downloading actions. We mitigated the incident by rolling back the change.

We are expanding test coverage and improving our deployment validation process to prevent recurrence of this issue in the future.

View →
GitHubresolvedMay 28, 2026 at 19:01 UTC

Disruption with OpenAI Models

May 28, 20:41 UTC
Resolved - On May 28th, 2026, between approximately 18:27 and 20:41 UTC, the GitHub Copilot service was degraded due to an issue with the Responses API of an upstream provider affecting the GPT-5.2, GPT-5.3-Codex, GPT-5.4, and GPT-5.5 models. Requests routed to these models via the Responses API returned elevated error rates, which also affected Copilot coding agent and Copilot code review. No other models were impacted.

We mitigated the incident by shifting traffic away from the affected models while the upstream provider deployed a fix.

GitHub is working to improve automated failover for the affected models and strengthen monitoring to prevent similar incidents in the future.

May 28, 20:06 UTC
Update - Open AI models are currently unavailable. We are shifting requests to other models to reduce impact.

May 28, 19:40 UTC
Update - We are investigating errors with Copilot requests using OpenAI models

May 28, 19:20 UTC
Update - Copilot is experiencing degraded performance. We are continuing to investigate.

May 28, 19:01 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedMay 28, 2026 at 16:00 UTC

Elevated Function Invocation Errors in Stockholm region (ARN1)

May 28, 16:59 UTC
Resolved - This incident has been resolved.

May 28, 16:37 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

May 28, 16:18 UTC
Identified - The issue has been identified and a fix is being implemented.

May 28, 16:00 UTC
Investigating - We've identified an issue where some customers may experience elevated error rates when invoking functions in the ARN1 Edge Region. We are currently investigating this issue.

View →
GitHubresolvedMay 28, 2026 at 01:13 UTC

Webhook APIs and UI Degraded

May 28, 01:32 UTC
Resolved - On May 28, 2026, between 00:54 UTC and 01:19 UTC, some users experienced errors when interacting with the Webhooks API, including webhook delivery history and configuration endpoints. On average, the error rate was 0.28% and peaked at 0.45%. This was due to a bug that caused a single Kubernetes pod to enter a CrashLoopBackOff after receiving a 500 with an empty response body from Cosmos DB.

We mitigated the incident by restarting the service. To prevent future incidents, we are pushing a change to handle this response scenario from Cosmos DB appropriately.

May 28, 01:27 UTC
Monitoring - The degradation affecting Webhooks has been mitigated. We are monitoring to ensure stability.

May 28, 01:13 UTC
Investigating - We are investigating reports of degraded performance for Webhooks

View →
GitHubresolvedMay 27, 2026 at 12:10 UTC

Incident with Pull Requests, Issues, Git Operations and API Requests

May 27, 13:16 UTC
Resolved - On May 27, 2026, between 12:07 UTC and 13:16 UTC, users experienced degraded performance for Git operations, Pull Requests, Issues, GraphQL API, and related services on github.com. During this time, operations that depended on Git file servers experienced elevated error rates (3.5% of pushes via HTTPS and 0.2% of pushes via SSH failed; no fetches/clones failed). An internal analytics component generated unexpectedly high load, which caused CPU saturation on the underlying infrastructure. This led to cascading slowdowns and errors across services that depend on Git operations. The issue was mitigated by stopping the offending component. Services began recovering shortly after mitigation and were fully restored by 13:16 UTC. We are taking steps to add resource limits and kill switches for internal analytics components to prevent similar issues in the future.

May 27, 12:54 UTC
Update - We're continuing to investigate degraded performance of Git operations, Issues and Pull requests.

May 27, 12:10 UTC
Investigating - We are investigating reports of degraded performance for API Requests, Git Operations, Issues and Pull Requests

View →
VercelresolvedMay 27, 2026 at 04:50 UTC

Elevated Errors on Vercel Dashboard (Project Overview Page)

May 27, 05:02 UTC
Resolved - This incident has been resolved.

We've rolled back a bad release to resolve the issue. Customers were unable to access the project overview page between 04:13 and 05:00 AM UTC.

May 27, 04:50 UTC
Investigating - We are currently investigating this issue.

View →
InngestMay 26, 2026 at 23:46 UTC

Downstream provider causing dashboard loading errors

Status: Resolved

The incident is now resolved.

Affected components
  • Inngest Dashboard (Operational)
View →
InngestMay 26, 2026 at 21:32 UTC

Issues connecting to data store for events/runs/metrics pages

Status: Resolved

The previous incident with the observability database has beens resolved, but there is a separate incident with a downstream provider.

Affected components
  • Inngest Dashboard (Operational)
View →
GitHubresolvedMay 26, 2026 at 15:44 UTC

Disruption with some GitHub services

May 26, 16:35 UTC
Resolved - On May 26, 2026, between 15:10 UTC and 16:35 UTC the Copilot service was degraded and many models were no longer available for use. On average, the error rate was ~5% and peaked at 11% of requests to the service. This was due to a change that introduced a configuration mismatch in HMAC signing credentials which caused the list of available models to be truncated. This was mitigated by rolling back the change. This rollback was complete by 15:34 UTC though users continued to see impact until cache TTLs expired.

We are working to improve our monitoring and error handling to reduce time to detection and better experience for issues like this in the future.

May 26, 16:24 UTC
Monitoring - The degradation affecting Copilot has been mitigated. We are monitoring to ensure stability.

May 26, 15:48 UTC
Update - Copilot is experiencing degraded performance. We are continuing to investigate.

May 26, 15:44 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendMay 26, 2026 at 15:36 UTC

Resend dashboard facing intermittent errors

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Webhooks (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Batch Emails (Operational)
  • Website (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • Broadcast Emails (Operational)
  • General API (Operational)
View →
GitHubresolvedMay 26, 2026 at 10:57 UTC

Incident with Actions and Pages

May 26, 13:18 UTC
Resolved - On May 26, 2026, between 10:40 UTC and 12:56 UTC, GitHub Actions jobs were degraded. From 10:40 to 12:16 UTC, all newly queued Actions runs failed to start. From 12:16 to 12:56 UTC, Actions runs that required downloading actions for their workflows continued to fail. GitHub Pages, Copilot Code Review, Copilot coding agent, Octoshift, and GitHub Enterprise Importer were also impacted due to their dependency on Actions.

This was caused by our automated account review system incorrectly suspending the service account used by GitHub Actions to authenticate workflow runs and download actions.

We mitigated by restoring the account at 12:16 UTC, marking it exempt from further automated review at 12:20 UTC, and redeploying a related service at 12:48 UTC to flush cached account state. Full recovery was confirmed at 12:56 UTC.

During this incident, a small number of Issues, PRs, Comments, and Discussions were marked as hidden when the service account was disabled. No data was lost. All content hidden because of this incident has been restored and full search index restoration is in progress.

To prevent a recurrence, we have added an allowlist of all service accounts that cannot be suspended by automated systems, and ensuring these protections are enforced consistently across all account management tooling. We are also improving diagnostic tooling for accounts and reducing cache propagation delays to shorten time to mitigate similar incidents in the future.

May 26, 13:01 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

May 26, 13:00 UTC
Monitoring - The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.

May 26, 12:37 UTC
Update - We have identified the cause of the authentication issues affecting GitHub Actions and are actively working on mitigation

May 26, 12:17 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

May 26, 11:53 UTC
Update - We are investigating authentication issues leading to failure in starting Actions runs and downloading actions. At this time the majority of Actions runs is impacted.

May 26, 11:19 UTC
Update - Actions is experiencing degraded availability. We are continuing to investigate.

May 26, 10:57 UTC
Investigating - We are investigating reports of degraded performance for Actions and Pages

View →
VercelresolvedMay 25, 2026 at 14:58 UTC

Delays Loading Runtime Logs

May 25, 16:22 UTC
Resolved - This incident has been resolved.

May 25, 16:05 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

May 25, 15:25 UTC
Identified - The issue has been identified and a fix is being implemented.

May 25, 14:58 UTC
Investigating - We are currently investigating elevated latency in loading runtime logs (Vercel Functions) in live mode. Log Drains are unaffected at this time.

View →
GitHubresolvedMay 25, 2026 at 09:02 UTC

Elevated rate of Git push errors

May 25, 09:02 UTC
Resolved - On May 25, 2026, between 09:02 UTC and 09:11 UTC, Git push operations over HTTPS and SSH experienced elevated failures. During this window, an average of 31% and a peak of 43% of push requests failed.

The incident was caused by a recently enabled code path that issued an unexpectedly expensive database query against a primary database. The resulting load exhausted the database's connection pool, which caused the push failures above. The acute impact resolved automatically as in-flight work completed. We mitigated the incident by disabling the feature flag controlling the new code path. To prevent recurrence, we have updated the affected background workflows to route reads to replica databases instead of the primary, removing the specific code pattern that caused this incident; broader follow-up work is underway to apply the same safeguard to similar workflows across GitHub.

View →
GitHubresolvedMay 23, 2026 at 16:01 UTC

Intermittent errors with app installation token authentication

May 23, 19:32 UTC
Resolved - On May 23, 2026 between 06:00 UTC and 19:12 UTC, GitHub experienced intermittent errors authenticating GitHub app installation tokens.

During this time, between 1-5% of app installation token authentication requests failed, with an average of 2.3% and the error rate peaking at approximately 5.4% around 14:00 UTC. Users may have experienced authentication failures when using GitHub Apps, including failures in Git operations and API calls using app installation tokens.

The issue was caused by an issue in a caching proxy component and was remediated by rolling back that component to a previous version. We are taking steps to improve monitoring for cache miss anomalies to ensure that token authentication remains functional during infrastructure changes and reviewing our protocol for testing and when we upgrade third-party dependencies.

May 23, 19:32 UTC
Update - This is fully mitigated, we will continue to monitor to ensure it does not reoccur.

May 23, 19:06 UTC
Update - We have identified and are applying additional mitigation and will continue to monitor for complete mitigation.

May 23, 18:42 UTC
Update - We see significant signs of mitigation and are monitoring for full mitigation.

May 23, 17:41 UTC
Update - We are seeing signs of mitigation and are continuing to monitor for complete mitigation. 

Next update in one hour.

May 23, 16:35 UTC
Update - We are continuing to investigate an elevated error rate of authentication failures for app installation tokens.  Next update in one hour.

May 23, 16:01 UTC
Update - We are seeing an increased rate of authentication failures for app installation tokens, affecting approximately 1% of tokens.  We are continuing to investigate.

May 23, 16:00 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedMay 23, 2026 at 10:43 UTC

Elevated Build Failures (GitHub connected projects)

May 23, 18:44 UTC
Resolved - This incident has been resolved.

Please refer to GitHub's status page post for more details: https://www.githubstatus.com/incidents/k5z4d1v1tqmt

May 23, 18:21 UTC
Monitoring - GitHub has implemented a fix, and we are monitoring the results.

May 23, 15:46 UTC
Identified - We have identified the issue and are working closely with GitHub to resolve it.

May 23, 14:30 UTC
Update - We are working closely with GitHub to resolve this issue. Failures are intermittent — if your deployment fails, redeploying should resolve it in the meantime. We will provide updates as the situation develops.

May 23, 12:15 UTC
Update - We are continuing to investigate this issue.

May 23, 11:09 UTC
Update - We've been observing increased failures of git operations with GitHub since around 06:00 UTC. Some deployments triggered by GitHub commits might have seen git-related errors in failed build logs.

CLI deployments are unaffected at this time.

May 23, 10:43 UTC
Investigating - We are currently investigating this issue.

View →
VercelresolvedMay 22, 2026 at 17:11 UTC

Build Failures for Some Next.js Deployments

May 22, 17:11 UTC
Resolved - Between 16:10 and 16:43 UTC on May 22, some customers using Next.js above 16.2.0-canary.28 with Preview Comments enabled experienced build failures during deployments. The issue has been mitigated and follow-up deployments should no longer encounter this error.

View →
InngestMay 21, 2026 at 20:20 UTC

Elevated latency in runs and event details lookups

Status: Resolved

The incident is now resolved and the system is fully operational.

Affected components
  • Inngest Dashboard (Operational)
  • API (REST and GraphQL) (Operational)
View →
VercelresolvedMay 21, 2026 at 15:01 UTC

Elevated Build Errors

May 21, 21:31 UTC
Resolved - This incident has been resolved.

May 21, 21:20 UTC
Monitoring - The mitigation was successfully rolled out, and builds are stable.

May 21, 19:55 UTC
Update - The root cause has been identified and we are rolling out a mitigation.

May 21, 15:01 UTC
Identified - We are currently investigating elevated build failures affecting a subset of Vite projects. Affected deployments may be timing out. We’ve identified an issue and are working on the fix.

View →
VercelresolvedMay 20, 2026 at 20:07 UTC

Missing Build CPU Minutes Usage Data

May 21, 06:08 UTC
Resolved - This incident has been resolved.

Usage data for Build CPU Minutes between May 15, 2026, 19:30, and May 20, 2026, 18:00 UTC is currently incomplete. Customers will see usage data for Build CPU Minutes catch up as we backfill the data for the affected window.

May 20, 20:07 UTC
Monitoring - We've identified an issue where some users may see missing Build CPU Minutes data on Usage pages in the Vercel Dashboard. The issue has been resolved, and we are backfilling the affected usage data.

View →
GitHubresolvedMay 20, 2026 at 16:58 UTC

Incident with Actions

May 20, 20:14 UTC
Resolved - On May 20, 2026, between 16:00 UTC and 17:45 UTC, GitHub Actions customers experienced run start delays exceeding 5 minutes. Approximately 4.5% of all runs were delayed during the impact window, with scale set jobs disproportionately affected. 30% of scale set jobs were delayed and 4% failed to start entirely.

The incident was caused by a misconfigured health check on an internal service that assigns jobs to runners. A brief latency spike in an upstream dependency triggered health check failures across several pods, removing them from service and concentrating load on the remaining capacity. The added load drove memory pressure that escalated into a cascading failure in one regional cluster, leaving it unable to self-recover.

Responders mitigated the incident by scaling capacity in the healthy regional clusters and draining traffic away from the impaired one, after which run start latency recovered. To prevent recurrence, we are strengthening our health check configuration to avoid cascading failure scenarios and evaluating automated mitigations to rebalance traffic when a region is degraded.

May 20, 19:41 UTC
Update - Customer impact has fully subsided. We are maintaining yellow status while we deploy a permanent fix to prevent recurrence.

May 20, 18:17 UTC
Update - We've applied a mitigation to fix the issues with queuing and running Actions jobs. We are seeing improvements in telemetry and are monitoring for full recovery.

May 20, 17:52 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

May 20, 17:46 UTC
Update - A subset of runners are taking longer than expected to connect, which may delay some jobs from beginning execution. We are actively working to mitigate the issue.

May 20, 16:58 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
ResendMay 20, 2026 at 15:13 UTC

Maintenance: Email Events - Scheduled Maintenance

Status: Complete

We've completed our Email Events maintenance, and they are now showing the up-to-date status.

Affected components
  • Email Events (Operational)
View →
GitHubresolvedMay 19, 2026 at 05:30 UTC

Incident with Copilot

May 19, 05:30 UTC
Resolved - On May 19, 2026, between 05:30 UTC and 14:50 UTC, some Copilot users experienced failures when using code completions, chat sessions, and cloud agent sessions. At peak impact, approximately 13% of Copilot API requests failed, and approximately 24% of remote sessions failed to initialize. A partial mitigation at 08:16 UTC reduced the Copilot API error rate to approximately 0.3%, but intermittent failures persisted until a full fix was deployed at 14:15 UTC and recovery was verified by 14:50 UTC.

The incident was caused by rate limits being exceeded on a shared infrastructure component. A recently enabled feature increased call volume to this component, and the combined load exceeded capacity limits as traffic increased during business hours.

We mitigated the incident by deploying a caching layer to reduce load on shared infrastructure. To prevent recurrence, we are separating rate limit scopes between services, adding monitoring for internal dependency rate limiting, and reducing redundant calls.

View →
VercelresolvedMay 18, 2026 at 21:42 UTC

Increased Function Invocation Errors - ERR_MODULE_NOT_FOUND

May 18, 23:52 UTC
Resolved - This incident has been resolved.

Some deployments created between May 18, 2026, 06:15 PM - 10:16 PM UTC may have experienced increased ERR_MODULE_NOT_FOUND function invocation errors due to a bad rollout. Deployments created outside of this window are unaffected.

The fix is being rolled out to existing deployments retrospectively and is expected to finish in the next few hours.

May 18, 22:19 UTC
Update - We are continuing to monitor for any further issues.

May 18, 22:16 UTC
Monitoring - A fix is rolled out for the function invocation errors affecting React Router 7 deployments. To recover, redeploy your application or use Instant Rollback to a previous deployment from the dashboard. We're continuing to monitor.

May 18, 22:11 UTC
Update - We're continuing to work on rolling out a fix. In the mean time, customers may use Instant Rollback to a previous deployment version as an immediate workaround in order to recover.

May 18, 21:53 UTC
Identified - The issue has been identified and a fix is being implemented.

May 18, 21:42 UTC
Investigating - We're investigating an issue where customers are currently experiencing application failures due to function invocation errors when using React Router 7.

View →
VercelresolvedMay 18, 2026 at 15:00 UTC

Support cases cannot be submitted

May 18, 15:41 UTC
Resolved - The incident is resolved and support cases can be submitted again.

May 18, 15:17 UTC
Monitoring - We have implemented a fix. New support cases can be created. We continue to monitor.

May 18, 15:00 UTC
Investigating - We are investigating an issue where support cannot be submitted through the dashboard.

View →
VercelresolvedMay 15, 2026 at 10:34 UTC

Workflow usage amounts are incorrectly calculated

May 15, 16:18 UTC
Resolved - The usage tracking system is fully recovered.

May 15, 15:39 UTC
Monitoring - The team fixed the incorrect calculations. We are resuming usage tracking and continuing to monitor.

May 15, 13:40 UTC
Update - We are continuing to work on the fix. Usage tracking for Workflow Storage remains paused. The incorrect calculations will not be charged.

May 15, 11:37 UTC
Identified - The team has identified the source of the incorrect usage data calculations and has paused usage tracking while a fix is being prepared.

May 15, 10:34 UTC
Investigating - Workflow Storage Retention and Workflow Storage Writes are being incorrectly calculated in usage data. The team is investigating and working to resolve the discrepancy.

View →
GitHubresolvedMay 15, 2026 at 08:14 UTC

Actions is experiencing degraded availability

May 15, 08:48 UTC
Resolved - On May 15, 2026, from approximately 07:43 UTC to 08:48 UTC, GitHub Actions experienced a degradation that caused workflow runs to fail or experience delayed starts for a subset of customers. The incident was triggered by a planned failover of supporting infrastructure used by GitHub Actions. During that operation, an automated service discovery update did not propagate correctly, which caused traffic to be routed incorrectly and increased request timeouts in a core dependency for workflow orchestration.

At peak impact, 42% of Actions runs failed. Downstream services that depend on Actions workflow execution were also impacted, including GitHub Pages and Copilot cloud services. At 08:12 UTC, responders manually corrected the service discovery routing issue. Timeout and failure rates recovered shortly after, and we continued monitoring until full stabilization was confirmed across all affected services. The incident was marked resolved at 08:48 UTC.

To prevent recurrence, we are implementing failover guardrails that validate service discovery state before completing failover operations, strengthening pre-flight and post-flight verification checks, and improving dependency resilience to reduce timeout cascades during infrastructure events.

May 15, 08:41 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

May 15, 08:29 UTC
Update - We are monitoring an issue that was affecting GitHub Actions and causing downstream issues in GitHub Coding Agent and GitHub Code Review Agent. The issue has resolved now but we are closely monitoring our systems for full recovery.

May 15, 08:27 UTC
Update - The degradation affecting Pages has been mitigated. We are monitoring to ensure stability.

May 15, 08:26 UTC
Update - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

May 15, 08:14 UTC
Update - Pages is experiencing degraded availability. We are continuing to investigate.

May 15, 08:13 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
GitHubresolvedMay 15, 2026 at 02:30 UTC

[Retroactive] Incident with GitHub.com

May 15, 02:30 UTC
Resolved - Beginning at 02:49 UTC on May 15 2026 and lasting until 03:04 UTC, GitHub.com was unavailable for a subset of customers. This impact has been mitigated and normal service resumed.

The issue was rooted in a sudden spike in traffic, with intermittent impact. We've identified the source of the traffic and prevented further disruption.

View →
GitHubresolvedMay 13, 2026 at 14:43 UTC

Incident with CodeQL

May 13, 16:03 UTC
Resolved - On May 13, 2026, between 14:31 and 16:03 UTC, the Code Scanning service experienced processing delays and 12% of check runs took over 15 minutes to complete. The delays were caused by replication lag due to an internal database migration, resulting in insufficient worker capacity for our high rate of job enqueues.

We mitigated the impact by scaling our processing workers by 34%. Code Scanning results returned to normal processing times after the mitigation was applied.

The capacity increases are permanent, and we are looking into more ways to decrease the load on our workers to help prevent this in the future.

May 13, 15:30 UTC
Update - CodeQL impact has been mitigated. We are continuing to monitor for durable recovery.

May 13, 15:26 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

May 13, 14:58 UTC
Update - We have applied a mitigation to increase processing capacity. We are continuing to monitor to confirm full recovery. We will provide another update by 15:30 UTC.

May 13, 14:43 UTC
Update - We are investigating delays affecting CodeQL, the code analysis engine used by Code Scanning. Some users may experience delayed or incomplete code scanning results. Our engineering team is investigating. We will provide another update by 15:15 UTC.

May 13, 14:41 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedMay 12, 2026 at 14:38 UTC

Incident with CodeQL, Webhooks, Notifications, and Slack Integration

May 12, 17:43 UTC
Resolved - On May 12, 2026, between 13:41 and 17:43 UTC, some services experienced delays in processing. For the Code Scanning service, 53% of check runs took over 15 minutes to complete. Additionally, notifications took an average of 22 minutes to be delivered and Slack integration webhooks took an average of 20 minutes to be delivered. The delays were caused by replication lag due to an internal database migration, resulting in insufficient worker capacity for our high rate of job enqueues.

We mitigated the impact by scaling our processing workers to handle the increased load. All services returned to normal processing times after the mitigation was applied.

We are working to create dedicated worker pools for some of our high usage shared queues to help prevent this in the future.

May 12, 17:43 UTC
Update - All services have fully recovered.

May 12, 16:59 UTC
Update - CodeQL has fully recovered. We're continuing to work on recovery for the remaining impacted services.

May 12, 16:29 UTC
Update - Webhooks have fully recovered. Continuing to work on recovery for the other services.

May 12, 16:28 UTC
Update - Webhooks is operating normally.

May 12, 16:18 UTC
Update - We've established that most delays are related to a queuing service and are working to scale out. Early signals from the scale-out are showing signs of recovery for some services. We'll provide an update when services are fully recovered.

May 12, 15:44 UTC
Update - Webhooks is experiencing degraded performance. We are continuing to investigate.

May 12, 15:42 UTC
Update - We're continuing to investigate issues with CodeQL actions workflows. We're additionally seeing delays for notifications, webhooks, and the Slack integration.

May 12, 15:13 UTC
Update - CodeQL actions are currently experiencing delays, which may result in those actions being stuck in a pending state or having failed due to a timeout.

May 12, 14:38 UTC
Investigating - We are investigating reports of degraded performance for CodeQL

View →
GitHubresolvedMay 11, 2026 at 14:25 UTC

Incident with high errors on Git Operations

May 11, 14:33 UTC
Resolved - On May 11th, 2026, between 14:00 UTC and 14:33 UTC, HTTP-based Git read operations were degraded. On average, the error rate was 2.8% and peaked at 7.5% of requests to the service. This was due to resource exhaustion in a networking gateway between GitHub.com’s frontend service for Git operations and a dependency service that performs authentication and authorization. Following the initial spike, the frontend service became stuck in a degraded state in one of our data centers, increasing time to mitigation.

We mitigated the incident by scaling the networking gateway and re-deploying the frontend service.

To reduce our time to detection and mitigation in the future, we are adding auto-scaling to the networking gateway, and resolving a bug which caused the frontend service to remain degraded.

May 11, 14:25 UTC
Investigating - We are investigating reports of degraded performance for Git Operations

View →
VercelresolvedMay 9, 2026 at 04:45 UTC

Vercel Queues, and Vercel Workflow runs were delayed

May 9, 05:40 UTC
Resolved - This incident has been resolved.

May 9, 04:45 UTC
Monitoring - Between 1:30 to 3:44 UTC, messages in Vercel Queues in iad1 were enqueued but not processed, and Vercel Workflows were blocked from making progress (i.e. remaining in pending / active states). This is recovering, Vercel Queues backlogs are being processed, and Vercel Workflows are unblocked.

View →
VercelresolvedMay 8, 2026 at 19:15 UTC

SSL Certificate Generation Delays

May 8, 19:15 UTC
Resolved - Between 18:36 and 19:04 UTC, issuing SSL certificates was delayed for new domains. These certificates have now been issued. Certificate renewals were not affected.

View →
ResendMay 8, 2026 at 17:47 UTC

Inbound emails with incorrect "To" field when forwarded

Status: Resolved

We have resolved the underlying issue by reverting the unintended breaking change and service has been resumed.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • Broadcast Emails (Operational)
  • Website (Operational)
  • SMTP (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Webhooks (Operational)
View →
VercelresolvedMay 8, 2026 at 16:33 UTC

Delays Processing Builds

May 8, 16:41 UTC
Resolved - This incident has been resolved.

May 8, 16:33 UTC
Monitoring - We have applied a fix and are observing recovery across all builds. We'll provide additional updates as-needed.

May 8, 16:30 UTC
Identified - We've identified an issue where some customers may experience delays in builds starting and/or builds stuck in an initializing state. We are applying a fix and will provide additional updates as they become available.

View →
VercelresolvedMay 8, 2026 at 01:18 UTC

Elevated errors across multiple services in IAD1 (Washington, D.C., USA)

May 8, 06:39 UTC
Resolved - We've deployed mitigations and all services are operating normally. Traffic may continue to reroute to nearby regions for the next few hours, but we will restore traffic gradually after verifying system availability in the IAD1 region.

May 8, 04:54 UTC
Monitoring - We implemented a fix and are monitoring the results. Backlogs for Workflows and Queues will be processed within the next few hours.

May 8, 02:33 UTC
Update - New workflow runs and queue messages are being processed now. Processing the backlog of existing workflow runs and queued messages is still being investigated.

May 8, 02:18 UTC
Update - We are continuing to work on a fix for this issue.

May 8, 02:13 UTC
Update - New messages to Vercel Workflows and Vercel Queues are being queued, but message processing is paused. Queued messages will be processed when service is restored.

May 8, 01:47 UTC
Update - Traffic to the IAD1 region has been re-routed to nearby regions. Functions configured in IAD1 will be invoked in a different region if failover regions are configured, but if no other regions are configured, the function will still be invoked in IAD1.

May 8, 01:18 UTC
Identified - Some Vercel functions that run in the IAD1 region are experiencing elevated invocation failures. We are investigating the issue and will share more information as it becomes available.

View →
VercelresolvedMay 7, 2026 at 21:14 UTC

Elevated Errors Creating New Deployments

May 7, 21:48 UTC
Resolved - This issue has been resolved and new deployments are being created successfully.

May 7, 21:14 UTC
Investigating - We are investigating an issue affecting new deployments that are stuck in the provisioning state. We will share more information as it becomes available.

View →
GitHubresolvedMay 7, 2026 at 05:02 UTC

CCR and CCA failing to start for PR comments

May 7, 06:56 UTC
Resolved - On May 7, 2026, between 04:12 UTC and 06:13 UTC, Copilot Cloud Agent and Copilot Code Review Agent sessions for pull requests were delayed or failed to start.

The issue was caused by follow-up recovery work from a separate Pull Requests incident (https://www.githubstatus.com/incidents/f5pb5d5mr9yh). As part of that recovery, we ran a large database migration, which caused replication delays on several replica hosts.

Although those replicas were not serving user traffic, our safeguards correctly treated the elevated replication lag as a signal to slow down writes to the affected database cluster. As a result, some pull request background processing was temporarily delayed. That processing is responsible for sending the internal events that Copilot agents use to begin work, so affected agents did not start until the database replicas caught up.

The system recovered once replication lag returned to normal and pull request processing resumed. We are reviewing how this safeguard interacts with recovery migrations so we can reduce the chance of similar secondary impact during future incident recovery work.

May 7, 06:14 UTC
Update - Copilot code review and cloud agents are starting again for pull requests, we are monitoring for full recovery.

May 7, 06:13 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

May 7, 05:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedMay 6, 2026 at 15:28 UTC

Incident with Pull Requests

May 6, 19:04 UTC
Resolved - On May 6, 2026 between 15:12 and 19:02 UTC creation of new pull request review threads on GitHub.com failed. This included new line comments and file comments on pull requests. Existing PRs and previously created comments were unaffected.

This incident was caused by a 32-bit integer key reaching its maximum value in a Vitess lookup table used during PR thread creation. The primary table had been migrated to a 64-bit integer key but the Vitesse lookup table remained 32-bit. Once the values in the primary table passed the available 32-bit ID space in the lookup table, attempts to create new review threads began failing, resulting in near 100% failure rate for new thread creation requests. We mitigated the issue by updating the impacted lookup table definitions across all shards to use 64-bit integer column types, increasing the available ID range and restoring normal operation. Service was fully restored once the schema changes competed globally.

To help prevent similar incidents, we are expanding existing monitoring of database columns to include Vitess lookup tables to notify in advance of any tables that is approaching a column size limit. This work is intended to provide earlier detection of columns approaching size limits before customer impact occurs.

May 6, 19:04 UTC
Update - Mitigations have been fully applied and we are seeing full recovery of functionality on Pull Request threads. We are continuing to monitor to ensure sustained recovery.

May 6, 17:52 UTC
Update - Creation of new Pull Request threads (including line and file comments) continues to be affected although we are seeing partial recovery.

A mitigation is being applied to continue to accelerate recovery with complete recovery expected by 8:00pm UTC.

Top-level comments on pull requests still function and should remain usable during recovery. Opening and merging pull requests, actions, and other pull request operations remain functional.

May 6, 16:20 UTC
Update - Creation of new Pull Request threads (including line and file comments) continues to be affected.

Top-level comments on pull requests still function and should remain usable during recovery. Opening and merging pull requests, actions, and other pull request operations remain functional.

A mitigation is being applied. Recovery is expected to be gradual, with complete recovery expected by 8:00pm UTC.

May 6, 16:07 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

May 6, 15:55 UTC
Update - Creation of new Pull Request threads (including line and file comments) continues to be affected. We have identified the cause of the issue and have started taking steps to mitigate this issue.

May 6, 15:28 UTC
Update - We are investigating failures for new thread creation on Pull Requests. Responses to existing pull request threads are unaffected.

May 6, 15:25 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
GitHubresolvedMay 6, 2026 at 11:21 UTC

Disruption with some GitHub services

May 6, 11:59 UTC
Resolved - On May 6, 2026 between 11:02 UTC and 11:13 UTC, users were unable to start or view Copilot Cloud Agent or remote sessions. During this time, requests to the session API returned errors, preventing users from creating new sessions or viewing existing ones. The issue was caused by a configuration change to the service's network routing that inadvertently removed the ingress path for the service. The team reverted the change at 11:13 UTC which restored service. The incident remained open until 11:59 UTC while the team verified full recovery. We are taking steps to improve our deployment validation process to prevent similar configuration changes from impacting production traffic in the future.

May 6, 11:59 UTC
Update - We have applied a mitigation and Copilot services have recovered.

May 6, 11:25 UTC
Update - We are investigating issues with the ability to start Copilot Cloud Agent sessions and view them.

May 6, 11:21 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendMay 6, 2026 at 10:48 UTC

Custom tracking domains failing to provision for new customers

Status: Resolved

This incident is resolved. All failed provisioning requests have been reprocessed, and custom tracking domain creation is operating normally.
View →
GitHubresolvedMay 6, 2026 at 07:19 UTC

Incident with Actions, we are investigating reports of degraded availability

May 6, 09:44 UTC
Resolved - On May 6, 2026, from approximately 06:45 UTC to 09:15 UTC, GitHub Actions Standard Ubuntu hosted runners were degraded. 17.1% of jobs requesting a standard runner failed.

This was caused by an unexpected data shape in the allocation configuration data for standard runners. That data was introduced as part of post-incident remediation work for an incident the previous day and caused new allocations to be blocked as load ramped up for the day. Removing that data at 08:51 allowed allocations to proceed and hosted runner pools to scale up and recover.

We are updating the filter logic for this allocation data to be resilient to abnormal data shapes and improving monitoring to alert when allocations are blocked, allowing the team to respond before customer impact starts.

May 6, 09:44 UTC
Update - Actions wait times have fully recovered.

May 6, 09:19 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

May 6, 09:08 UTC
Update - We've applied a mitigation to fix the issues with queuing and running Actions jobs. We are seeing improvements in telemetry and are monitoring for full recovery.

May 6, 08:00 UTC
Update - Actions is experiencing issues with ubuntu standard hosted runners leading to high wait times. We are actively investigating the issue

May 6, 07:19 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
VercelresolvedMay 6, 2026 at 00:10 UTC

Increased error rate in Workflows

May 6, 00:19 UTC
Resolved - This incident has been resolved.

May 6, 00:13 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

May 5, 23:30 UTC
Identified - We are experiencing increased error rates with Workflows running on Vercel. We have identified the issue and are implementing a fix.

View →
ResendMay 5, 2026 at 17:01 UTC

Maintenance: Logs Page - Scheduled Maintenance

Status: Complete

We've successfully completed our maintenance, and the logs page is now behaving normally.

Affected components
  • Dashboard (Operational)
View →
GitHubresolvedMay 5, 2026 at 16:49 UTC

Increased Latency and Failures for SSH Git Operations

May 5, 18:35 UTC
Resolved - Between approximately 14:00 and 16:10 UTC on May 5, 2026, SSH-based Git operations experienced elevated latency and intermittent failures. On average, the error rate was 0.46% and peaked at 0.6% of SSH write requests. HTTP-based Git operations, including web UI and HTTPS clones, were not affected.

The impact was caused by reduced SSH capacity at one of our data center sites. During a period of high traffic, the remaining hosts became overloaded, leading to connection exhaustion and some failures for SSH-based operations.

Additional capacity was provisioned to expand SSH capacity and resolve the incident. The expanded capacity was fully online by 18:18 UTC.

To reduce the likelihood of similar incidents, we will implement faster scaling solutions for SSH infrastructure and improved alerting for host availability and capacity thresholds.

May 5, 18:35 UTC
Update - We've completed our mitigation to prevent further impact. At this time the incident is considered resolved.

May 5, 18:25 UTC
Monitoring - The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.

May 5, 17:26 UTC
Update - We're continuing to work on preventing further impact from the earlier issue. No SSH-based impact is expected at this time. We'll post new updates if impact recurs or once our mitigation is in place.

May 5, 17:23 UTC
Investigating - Git Operations is experiencing degraded performance. We are continuing to investigate.

May 5, 16:54 UTC
Monitoring - Between approximately 14:00 and 16:10 UTC, customers using SSH-based Git operations may have experienced elevated latency and failures. HTTP-based operations were not impacted. We've identified a suspected root cause and are working to implement a mitigation to prevent further impact.

May 5, 16:49 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedMay 5, 2026 at 15:15 UTC

Failures to load data across multiple services on the Vercel dashboard, API, and CLI

May 6, 00:38 UTC
Resolved - All services remain operational. New data is being processed as expected. The majority of the backfill is complete and a long tail of data will be processed within the next couple of hours.

May 5, 17:14 UTC
Update - All systems are operational. We are continuing to monitor the backfill process, which is expected to take a couple of hours. New data is being processed as expected.

May 5, 17:12 UTC
Update - We are continuing to monitor for any further issues.

May 5, 16:50 UTC
Monitoring - Impacted services have recovered. We are continuing to monitoring, while backfilling remaining missing data.

May 5, 16:37 UTC
Update - Reading data from impacted services has recovered. We are working to backfill any missing data in Observability.

May 5, 16:30 UTC
Update - We are seeing services start to recover as we incrementally apply the fix.

May 5, 15:58 UTC
Update - We are continuing to work on a fix for this issue.

May 5, 15:15 UTC
Identified - We've identified an issue where various services in the Vercel dashboard, API, and CLI are failing to load data. These services include Observability, Speed Insights, Alerts, Usage, Web Analytics, Firewall, and CDN observability. Broadly, areas backed by analytical data are affected.

We have identified the issue and are working on a fix.

Production applications running on Vercel are not impacted.

View →
GitHubresolvedMay 5, 2026 at 13:37 UTC

Incident with Actions

May 5, 17:26 UTC
Resolved - On May 5, 2026, from approximately 13:22 UTC to 17:05 UTC, GitHub Actions hosted runners in the East US region were degraded. 13.5% of jobs requesting a standard runner failed and ~16% of requested Larger Runners with private networking pinned to East US failed or were delayed by more than 5 minutes. Copilot Code Review requests were also impacted. Approximately 8,500 code review requests timed out during this window. Affected users saw an error comment on their pull requests and were able to retry by re-requesting a review. Most runner requests were picked up by other regions automatically, but a portion of requests still routing to East US were impacted.

This was triggered by a scale-up operation for hosted runner VMs in the East US region. This is a regular operation, but the VM create load hit an internal rate limit when VM creates pull images from storage. Existing backoff logic was not triggered because of the response code returned in this case. The rate limiting and VM creation failures were mitigated by reducing load to allow for recovery and allowing queued work to be processed. By 15:34 UTC, queued and failed job assignments were mostly mitigated, with less than 0.5% of runner assignments impacted between 15:34 and full recovery at 17:05.

We are improving our system’s throttling behavior when limits occur, improving our controls to more quickly mitigate similar situations in the future, and reviewing all limits end-to-end for similar operations. We also immediately paused all scale and similar operations until these changes are in place and validated.

May 5, 17:11 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

May 5, 17:11 UTC
Update - Standard hosted runners have now reached full recovery. Hosted Runners with Private Networking in the East US region remain degraded as we continue working with our compute provider to restore capacity. Hosted Runners with private networking can fail over to a different Region to mitigate the issue.

May 5, 16:33 UTC
Update - We've seen signs of recovery for Standard Hosted Runners and are continuing to monitor for full recovery. Hosted Runners with Private Networking in the East US region remain affected as we continue working with our compute provider to restore capacity.

May 5, 15:54 UTC
Update - We've applied a mitigation for long queue times and failures on Standard Hosted Runners and are monitoring for full recovery. Hosted Runners with Private Networking in the East US region remain affected as we continue working with our compute provider to restore capacity.

May 5, 15:12 UTC
Update - We are working with our compute provider to alleviate elevated queue times and failures for Actions Jobs running on Hosted Runners in the East US region affecting 10% of runs. Hosted Runners with private networking can fail over to a different Region to mitigate the issue.

May 5, 14:14 UTC
Update - We are investigating elevated queue times and failures on Actions Jobs running on Hosted Runners in East US affecting 8% of runs. Hosted Runners with private networking can fail over to a different Azure region to mitigate the issue.

May 5, 13:48 UTC
Update - We are investigating elevated queue times on Actions Jobs running on Standard Hosted Runners in East US affecting 10% of runs

May 5, 13:37 UTC
Investigating - We are investigating reports of degraded availability for Actions

View →
VercelresolvedMay 4, 2026 at 21:30 UTC

Request Failures in ICN1 (Seoul, South Korea)

May 4, 21:30 UTC
Resolved - Between 21:45–21:52 UTC, some users may have experienced request failures in ICN1. The issue has been identified, a fix has been applied, and the issue has been resolved.

View →
GitHubresolvedMay 4, 2026 at 15:48 UTC

Incident with Issues and Webhooks

May 4, 16:40 UTC
Resolved - On 2026-05-04 at 3:37:17 PM UTC we detected increased latency on issues resulting in timeouts, and elevated 500 errors on webhooks. A scheduled workload drove high utilization on the primary host of a critical datastore, saturating the connection pool. We paused the job to mitigate the problem at 4:40:05 PM UTC and have implemented measures to prevent recurrence.

May 4, 16:36 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

May 4, 16:35 UTC
Update - Webhooks is operating normally.

May 4, 16:35 UTC
Update - The degradation affecting Codespaces has been mitigated. We are monitoring to ensure stability.

May 4, 16:34 UTC
Update - The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

May 4, 16:32 UTC
Update - Pull Requests is operating normally.

May 4, 16:29 UTC
Update - Pages is operating normally.

May 4, 16:29 UTC
Update - Latency across services has normalized. We are continuing to investigate the root cause and prevent reoccurrence.

May 4, 16:28 UTC
Update - Actions and Packages are operating normally.

May 4, 16:25 UTC
Update - Git Operations is operating normally.

May 4, 16:06 UTC
Update - Pages is experiencing degraded performance. We are continuing to investigate.

May 4, 16:05 UTC
Update - Codespaces is experiencing degraded performance. We are continuing to investigate.

May 4, 15:56 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

May 4, 15:51 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

May 4, 15:51 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

May 4, 15:50 UTC
Update - Packages is experiencing degraded performance. We are continuing to investigate.

May 4, 15:48 UTC
Update - Git Operations is experiencing degraded performance. We are continuing to investigate.

May 4, 15:48 UTC
Update - We are investigating Increased latency and timeouts across multiple GitHub services.

May 4, 15:45 UTC
Investigating - We are investigating reports of degraded performance for Issues and Webhooks

View →
VercelresolvedMay 1, 2026 at 11:30 UTC

Elevated Functions Invocation Errors in ICN1 (Seoul, South Korea) Region

May 1, 11:30 UTC
Resolved - Between May 1, 11:01 am – May 1, 11:09 am UTC, users may have seen increased function invocation errors for traffic originated near the ICN1 Vercel CDN region. Requests to static assets/cached content were unaffected this time. We have reverted the change that caused the issue to mitigate this.

View →
VercelresolvedApr 30, 2026 at 18:57 UTC

Elevated Build Errors

Apr 30, 19:45 UTC
Resolved - This incident has been resolved.

Apr 30, 19:36 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 30, 19:22 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 30, 18:57 UTC
Investigating - We are currently investigating this issue.

View →
VercelresolvedApr 30, 2026 at 17:30 UTC

Elevated Build Errors for Secure Compute/Static IPs Projects

Apr 30, 18:00 UTC
Resolved - This incident has been resolved.

Apr 30, 17:49 UTC
Update - We are continuing to monitor for any further issues.

Apr 30, 17:49 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 30, 17:39 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 30, 17:30 UTC
Investigating - We are currently investigating this issue.

View →
VercelresolvedApr 29, 2026 at 15:28 UTC

Elevated Domains Errors

Apr 29, 16:42 UTC
Resolved - This incident has been resolved.

Apr 29, 16:20 UTC
Monitoring - A fix has been applied and we are seeing signs of recovery across domains operations. We are continuing to monitor and will provide additional updates as needed.

Apr 29, 15:28 UTC
Investigating - We are investigating an issue causing elevated error rates across domains operations, including renewals, nameserver changes, search, and checkout. We will provide additional updates as they become available.

View →
InngestApr 28, 2026 at 22:48 UTC

Increased function execution latency

Status: Resolved

Performance on all shards is back to normal levels. The degradation was caused by a change that aimed to improve concurrency metrics for Inngest accounts. The change, while out for a full 24 hours and fairly benign, began to compound and produced slowness for some queue shards within our system earlier today. We have reverted that change and are working to understand the performance impact of this change.

Affected components
  • Function execution (Operational)
View →
ResendApr 28, 2026 at 20:01 UTC

Email Events Displaying Queued

Status: Resolved

This incident is resolved. Webhook deliveries and metrics page updates have been stable since our fix was deployed. Sending and receiving were unaffected throughout. Thanks for your patience.

Affected components
  • Email Events (Operational)
View →
VercelresolvedApr 28, 2026 at 15:46 UTC

Errors Deploying Templates

Apr 28, 16:53 UTC
Resolved - This incident has been resolved.

Apr 28, 16:44 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 28, 16:15 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 28, 15:46 UTC
Investigating - We're investigating reports of unexpected errors when trying to create and deploy new projects from templates. We will provide additional updates as they become available.

View →
GitHubresolvedApr 28, 2026 at 14:17 UTC

Incomplete pull request results in repositories

May 1, 04:15 UTC
Resolved - On April 28, 2026, at approximately 14:07 UTC, GitHub received reports that pull requests were missing from search results across global and repository /pulls pages.

The issue was caused by a manually invoked repair job intended for a single repository, which was executed without the required safety flags. During execution of the repair job, the database query remained correctly scoped to the repo’s PR IDs. However, the Elasticsearch reconciliation logic did not apply the same scope. It interpreted the min and max PR IDs as a continuous range, causing unrelated PR documents across other repos to be marked for deletion. This resulted in the removal of 1,789,756,838 PR documents from the search index, approximately 49% of indexed PR documents.

Customer impact was limited to PR search and list discoverability. Primary storage was unaffected, and there was no impact to opening, updating, or merging PRs.

The issue was identified ~10 minutes after initial customer reports. Because it affected search index completeness rather than service availability, it was not caught by existing monitoring.

The root cause was a flaw in the search document repair framework: it allowed a scoped reconciliation to run without enforcing a matching Elasticsearch query scope. This created a destructive mismatch between the source-of-truth and the index. The issue was compounded by the ability to trigger the job from the production console without safety defaults. Prior testing focused only on safe backfill scenarios and did not cover this reconciliation path. Additionally, there was no automated detection for large-volume deletions in Elasticsearch.

We mitigated the incident through three parallel actions: (1) Deployed a MySQL-backed search fallback for the most active repos by traffic to restore PR visibility for highly impacted users (2) Initiated a snapshot restore and reindex process to repopulate missing pull request documents in Elasticsearch (3) Added a degradation notice on PR pages to inform users of incomplete search results while recovery was in progress. The incident was resolved on May 1, 2026 at 4:15 UTC, following completion and validation of the reindex process.

To prevent recurrence, we are prioritizing improvements to the repair framework and safeguards. These include enforcing scoped query alignment between primary storage and Elasticsearch, preventing destructive operations without explicit opt-in, strengthening guardrails for manual repair jobs, and evaluating restrictions on production console access.

In parallel, we are expanding automated test coverage for reconciliation safety invariants and introducing detection for anomalous deletion patterns in Elasticsearch so similar issues can be identified or blocked earlier.

We are committed to improving the safety and reliability of our repair systems and ensuring that operational workflows are resilient to both software defects and manual invocation risks.

May 1, 04:11 UTC
Update - This incident has been resolved. Search and indexing functionality for pull requests are now fully restored. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Apr 30, 03:49 UTC
Update - We have repaired the missing search records for affected Pull Requests and are working to identify and repair records left in a stale state after the recovery.

Apr 29, 22:22 UTC
Update - We have restored search/indexing functionality for over 99% of impacted pull requests. We are continuing to address the remaining affected pull requests and are reviewing outstanding gaps as part of the restoration process.

Apr 29, 00:40 UTC
Update - Mitigation is in progress, with full recovery of impacted pull request listings expected within approximately 24 hours.

Apr 28, 22:46 UTC
Update - We have made an interim mitigation to improve availability for some impacted repositories while reindexing continues, and we are actively monitoring the indexing progress.

Apr 28, 21:43 UTC
Update - Elastic search reindexing of pull requests is continuing. All data is preserved, but may not be available on pages relying on elasticsearch until the reindex is complete.

Pages and APIs that do not rely on elasticsearch, including the GitHub CLI (gh pr list) and API (/repos/{owner}/{repo}/pulls), are not impacted and can be used to retrieve pull request data in the interim.

Apr 28, 15:58 UTC
Update - We are actively reindexing the remaining ElasticSearch indexes. Our priority is ensuring correctness and avoiding further impact.  We are taking a measured approach to safely backfill data and will share additional updates as progress continues.

Apr 28, 14:51 UTC
Update - After yesterday’s incident, we are investigating cases where /pulls and /repo/pulls pages are not showing all indexed pull requests. This is because our Elasticsearch cluster does not currently contain all indexed documents.

No pull request data has been lost. As pull requests are updated, they will be reindexed. We are also working on accelerating a full reindex so these pages return complete results again.

Apr 28, 14:17 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
GitHubresolvedApr 28, 2026 at 13:59 UTC

Disruption with some GitHub services

Apr 28, 17:09 UTC
Resolved - On April 28, 2026, from approximately 12:41 UTC to 17:09 UTC, GitHub Actions jobs using Standard Ubuntu 22 and Ubuntu 24 hosted runners experienced run start delays. Approximately 8% of hosted runner jobs using Ubuntu 22 and Ubuntu 24 experienced delays greater than 5 minutes or failures. Larger and self-hosted runners were not impacted.

This was caused by a performance regression introduced in the VM reimage process. That reimage delay lowered the overall capacity of runners available to pick up new jobs. This was mitigated with a rollback to a known good image version.

We are addressing the core issue with reimage performance and improving the granularity of reimage telemetry across our services and our compute provider to more quickly diagnose similar issues in the future. Finally, we are evaluating other rollout changes to automatically detect similar regressions.

Apr 28, 17:08 UTC
Monitoring - Actions is operating normally.

Apr 28, 16:36 UTC
Update - Less than 1% of hosted ubuntu-latest runs are delayed. We’re working through remaining steps to restore runner capacity.

Apr 28, 15:41 UTC
Update - Currently less than 2% of hosted ubuntu-latest and ubuntu-24.04 runs are delayed or failing. We are continuing to monitor for full recovery.

Apr 28, 15:20 UTC
Update - We've applied a mitigation to unblock running Actions. We're continuing to monitor.

Apr 28, 14:49 UTC
Update - We're still investigating the root cause for run start delays and failures for Actions hosted Ubuntu jobs, around 5% of jobs are impacted as of now.

Apr 28, 14:02 UTC
Investigating - Actions is experiencing degraded performance. We are continuing to investigate.

Apr 28, 13:59 UTC
Monitoring - Actions is experiencing capacity constraints with hosted ubuntu-latest and ubuntu-24.02, leading to high wait times. Other hosted labels and self-hosted runners are not impacted.

Apr 28, 13:59 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 28, 2026 at 09:00 UTC

Delayed Usage Data for Observability Events

Apr 28, 09:00 UTC
Resolved - We have identified an issue resulting in delayed and/or missing usage data for observability events between 9:00 - 10:30 UTC and 16:50 - 19:30 UTC.

The issue has been identified and resolved. Usage data for observability events for the affected time buckets have been recovered. The affected windows will continue to show missing data - the recovered data will show up as additional usage at recovery timestamps. The total usage count for the period is accurate and no additional actions are required for customers at this time.

View →
GitHubresolvedApr 27, 2026 at 22:46 UTC

GitHub search is degraded

Apr 27, 22:46 UTC
Resolved - On April 27, 2026 between 16:15 UTC and 22:46 UTC, GitHub search services experienced degraded connectivity due to saturation of the load balancing tier deployed in front of our search infrastructure. This resulted in intermittent failures for services relying on our search data including Issues, Pull Requests, Projects, Repositories, Actions, Package Registry and Dependabot Alerts. The impact was varied by search target, with services seeing up to 65% of searches timing out or returning an error between 16:15 UTC and 18:00 UTC.

We detected the drop in search results through our ongoing monitoring and declared an incident at 16:21 UTC when we determined the issues would not self-heal. We tracked the incident as mitigated as of 21:33 UTC and monitored the systems until 22:46 UTC when we declared the incident resolved. Our existing monitoring did not classify the increased scraping as a risk and this dimension of the incident was only discovered while working to mitigate.

The saturation was caused by a large influx of anonymous distributed scraping traffic that was crafted to avoid our public API rate limits. This scraping traffic made up 30% of the day’s total search traffic, but it was concentrated within a four-hour period. The traffic originated from over 600,000 Unique IP addresses, with matching actor information across the board.

To mitigate, we immediately focused on relieving pressure from the load balancers while simultaneously working on scaling the load balancing tier, blocking the anomalous traffic and applying tuning to the balancers to fully resolve the incident.

Looking ahead, we’ve not only scaled the load balancer tier, but applied optimizations to improve our connection handling and re-use to reduce the possibility that a saturation event like this can re-occur. We’ve also added new monitors and controls within the platform to allow us to restrict anonymous traffic to mitigate the impact to our registered users.

Apr 27, 22:44 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 27, 22:35 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 27, 21:33 UTC
Monitoring - The degradation affecting Actions, Issues, Packages and Pull Requests has been mitigated. We are monitoring to ensure stability.

Apr 27, 21:32 UTC
Update - We've applied a mitigation and continuing to monitor

Apr 27, 20:06 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Apr 27, 19:50 UTC
Update - We have identified the source of the additional load causing stress on our ElasticSearch clusters. We have disabled the source of that load and are seeing signs of recovery

Apr 27, 18:19 UTC
Update - Pull Requests is experiencing degraded availability. We are continuing to investigate.

Apr 27, 18:17 UTC
Update - We're continuing to see connectivity issues reaching elasticsearch. Impact on downstream services will be intermittent as we find the root cause

Apr 27, 17:35 UTC
Update - Users are experiencing intermittent failures to view issues, pull requests, projects and Actions workflow runs.

We are still investigating and attempting mitigations. We will provide further updates.

Apr 27, 16:53 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Apr 27, 16:39 UTC
Update - Packages is experiencing degraded performance. We are continuing to investigate.

Apr 27, 16:36 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Apr 27, 16:33 UTC
Update - Customers across GitHub are experiencing failures with searches. Examples include: workflow run failures, projects failing to load, and timed out search requests. This is due to an ongoing infrastructure issue that we have been investigating.

Apr 27, 16:31 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
GitHubresolvedApr 27, 2026 at 19:02 UTC

Disruption with some GitHub services

Apr 27, 19:02 UTC
Resolved - On April 22, 2026 from 18:49 to 19:32 UTC , the Copilot Cloud Agent service began failing during session execution for users running the Agent HQ Codex agent. Codex agent sessions failed to start for all entry points (issue assignment, @copilot comment mentions). 0.5% of total Copilot Cloud Agent jobs were impacted (~2,000 failed jobs). Copilot and other agent sessions were unaffected.

This was caused by a model resolution mismatch in Codex agent sessions, resulting in an incompatible model being used at runtime. A mitigation was deployed to select a stable default model for Codex agent sessions.

We are working to harden the underlying model-resolution path so it correctly scopes to the requesting agent's supported models to prevent similar failure mode in the future.

Apr 27, 19:01 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 27, 17:26 UTC
Update - We've found the issue and are working on deploying a solution to get Codex agent runs working again.

Apr 27, 17:01 UTC
Update - Copilot Cloud Agent (CCA) jobs using the Codex agent are failing after starting. To avoid this issue, please choose a different agent. We are investigating the cause and working towards remediation

Apr 27, 16:48 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendApr 27, 2026 at 17:33 UTC

Email Events Displaying Queued

Status: Resolved

This incident is resolved. Email events are updating in real time and have been stable since our fix was deployed. Sending and delivery were unaffected throughout. Thank you for your patience.

Affected components
  • Email Events (Operational)
View →
InngestApr 25, 2026 at 18:04 UTC

Degraded performance for REST API (runs/events data)

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • API (REST and GraphQL) (Operational)
View →
VercelresolvedApr 25, 2026 at 04:19 UTC

Errors Purchasing Domains, AI Gateway Credits

Apr 25, 04:19 UTC
Resolved - This incident has been resolved.

Apr 25, 04:12 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 25, 03:07 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 25, 02:20 UTC
Update - We're continuing to investigate reports of customers being unable to checkout with domains purchases. We're also investigating reports of being unable to purchase AI gateway credits. We'll provide additional updates as they become available.

Apr 25, 00:39 UTC
Investigating - We're investigating reports of customers being unable to checkout with domains purchases. We'll provide additional updates as they become available.

View →
GitHubresolvedApr 25, 2026 at 00:36 UTC

Delays with Actions Jobs for Larger Runners using VNet Injection in the East US region

Apr 25, 00:36 UTC
Resolved - On April 24, 2026, from approximately 11:39 UTC to April 25, 2026 at 00:15 UTC, GitHub Actions experienced delays and timeouts for Larger Hosted Runner jobs using VNet injection in the East US region without a failover region configured. Standard and Self-hosted runners were not impacted. This was caused by backend failures in our compute provider’s provisioning, scaling, and update operations for VMs in the East US region and mitigated by a rollback across all affected Availability Zones. More detail is available at https://azure.status.microsoft/en-us/status/history/?trackingId=5GP8-W0G.

We are working to improve the reliability of our annotations for jobs impacted by regional issues and are adding system log notifications as an additional customer communication channel alongside annotations.

VNet Failover is also now in public preview, allowing customers to evacuate Larger Hosted Runners using VNet injection in cases like this.

Apr 24, 19:14 UTC
Update - This is related to the public impact, "Multiservice impact for Azure Workloads in East US" shared at https://azure.status.microsoft/

Apr 24, 19:09 UTC
Update - We are investigating reports of degraded performance for Larger Runners with vnet injection in East US and we are working with our service provider on mitigation.

Apr 24, 19:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendApr 24, 2026 at 19:26 UTC

Email events delayed

Status: Resolved

The processing backlog is fully cleared. All systems are operational.

Affected components
  • Email Events (Operational)
View →
GitHubresolvedApr 23, 2026 at 19:50 UTC

Incident with Pull Requests

Apr 23, 21:43 UTC
Resolved - On April 23, 2026, between 16:05 UTC and 20:43 UTC, the Pull Requests service experienced a regression affecting merge queue operations. PRs merged via merge queue using the squash merge method produced incorrect merge commits when the merge group contained more than one PR. In affected cases, changes from previously merged PRs and prior commits were inadvertently reverted by subsequent merges.

During the impact window 2,092 pull requests were affected. The issue did not affect pull requests merged outside of merge queue, nor merge queue groups using the merge or rebase methods.
It took approximately 3 hours and 33 minutes to identify the issue. The change completed deployment at approximately 16:05 UTC, and we became aware at 19:38 UTC following an increase in customer support inquiries. Because the issue affected merge commit correctness rather than availability, it was not detected by existing automated monitoring and was identified through customer reports.

The regression was introduced by a new code path that adjusted merge base computation for merge queue ref updates. This code path was intended to be gated behind a feature flag for an unreleased feature, but the gating was incomplete.

As a result, the new behavior was inadvertently applied to squash merge groups, producing an incorrect three-way merge. This caused subsequent squash merges to revert changes from earlier pull requests and, in some cases, changes between their starting points.

We mitigated the incident by reverting the code change and force-deploying the fix across all environments. After resolution, we identified affected repositories and sent targeted remediation instructions to repository administrators with step-by-step recovery guidance.

The regression was not identified during internal validation. Existing test coverage primarily exercised single-PR merge queue groups, which did not exhibit the faulty base-reference calculation. Because automated checks did not validate merge correctness for multi-PR squash groups, the defect surfaced only in production.

To prevent recurrence, GitHub is expanding test coverage for merge correctness validation. We are broadening automated coverage for merge queue operations, including regression checks that validate resulting Git contents across supported configurations, so issues affecting merge correctness are caught before reaching production.

We are committed to ensuring the correctness and reliability of merge queue operations. These actions will reduce the risk of similar regressions and improve confidence in future changes to the Pull Requests service.

Apr 23, 21:18 UTC
Update - We have resolved a regression present when using merge queue with either squash merges or rebases. If you use merge queue in this configuration, some pull requests may have been merged incorrectly between 2026-04-23 16:05-20:43 UTC.

This behavior is still present in GitHub Enterprise Cloud with Data Residency, and we are rolling out the same fix.

Apr 23, 20:47 UTC
Update - Pull Requests is operating normally.

Apr 23, 19:58 UTC
Update - We have identified a regression in merge queue behavior present when squash merging or rebasing. We have identified the root-cause and are in the process of reverting the change.

Apr 23, 19:50 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
GitHubresolvedApr 23, 2026 at 19:28 UTC

Disruption with users unable to start Claude and Codex agent task from the web

Apr 23, 19:42 UTC
Resolved - Between 18:45 and 19:42 UTC on April 23, users were unable to start new agent tasks using either Claude or Codex agent on github.com. This was caused by a code change to how Copilot mission control routes task creation requests. Ongoing agent tasks and other Copilot agent features were not affected. We mitigated the impact by reverting the breaking change. We are adding extra monitoring and integration test coverage for the task creation path to prevent future recurrence.

Apr 23, 19:33 UTC
Update - We have identified the root cause of the issue and are working on mitigation.

Apr 23, 19:28 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendApr 23, 2026 at 18:36 UTC

Degraded performance in email sending via API and SMTP

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Email Events (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Dashboard (Operational)
View →
GitHubresolvedApr 23, 2026 at 16:52 UTC

Incident with multiple GitHub services

Apr 23, 17:30 UTC
Resolved - On April 23, 2026, between 16:03 UTC and 17:27 UTC, multiple GitHub services experienced elevated error rates and degraded performance due to DNS resolution failures originating from our DNS infrastructure in our VA3 datacenter. Approximately 5–7% of overall traffic was affected during the impact window:

- Webhooks: ~0.35% of API requests returned 5xx (peak ~0.39%). ~0.88% of requests exceeded 3s latency; at peak, >3s responses represented ~10% of Webhooks API traffic.

- Copilot Metrics: ~9% of Copilot Insights dashboard requests returned 5xx.

- Copilot cloud agents: ~10% of cloud agent sessions were affected and failing.

- Octoshift: 0.88% of active repo migrations failed and 79% saw elevated durations (avg. 5.2 min) during this period.

- Git Operations: averaged 1.25% errors over the duration of the incident, with a peak of 2.07% errors.

- Actions: Workflow run status updates experienced delays of up to ~8s over the duration of the incident window.

Our DNS infrastructure in VA3 entered a degraded state and began intermittently returning NXDOMAIN responses and timing out on lookups for both internal service discovery and external endpoints. This caused a cascading impact across the dependent services listed above.

We identified a specific load pattern under which our DNS resolvers began failing. The evidence points to a recently introduced traffic-balancing mechanism, rolled out progressively to support our growth, as the root cause. We have since reverted this change.

We are immediately prioritizing investments in a more controlled rollout and validation process, including a dedicated environment to safely shadow production DNS traffic and detect these failure modes before they can affect production.

Apr 23, 17:10 UTC
Update - Webhooks is operating normally.

Apr 23, 17:04 UTC
Update - Many services are mitigated and are validating the remaining services.

Apr 23, 17:03 UTC
Update - The degradation affecting Actions and Copilot has been mitigated. We are monitoring to ensure stability.

Apr 23, 16:52 UTC
Update - We have identified the root problem and are working on mitigation.

Apr 23, 16:34 UTC
Update - Actions is experiencing degraded performance. We are continuing to investigate.

Apr 23, 16:19 UTC
Update - We are investigating multiple unavailable services.

Apr 23, 16:12 UTC
Investigating - We are investigating reports of degraded availability for Copilot and Webhooks

View →
GitHubresolvedApr 23, 2026 at 15:18 UTC

Investigating errors on GitHub

Apr 23, 15:18 UTC
Resolved - On April 23, 2026 between 14:30 UTC and 15:18 UTC multiple services were degraded on github.com. During this time approximately 1.5% of all web requests resulted in a 5xx status and unicorn pages for github.com users. We also saw elevated error rates across Actions workflow runs, Copilot, Codespaces and Packages, leading to degraded experiences during this timeframe. Codespaces impact peaked at 45% failures for create requests and 65% failures for resume requests. Packages impact was mainly Maven related with 50% failure rates in downloads and 70% failure rates in uploads. Actions experienced a peak of 8% of failed jobs and up to 85% of jobs impacted by run start delays of more than 5 minutes.

This was due to a configuration change to an internal billing service that led to a cache being overwhelmed and causing requests to time out. These timeouts cascaded across multiple services and eventually caused requests to queue up and exhaust web request workers.

This configuration change was reverted at 14:42 UTC and following this, all services began to see recovery immediately.

To prevent this situation in the future, we are taking steps to ensure that failures and timeouts in the billing service don’t cascade to other services causing impact. This includes implementing more aggressive timeouts on callers of these billing services, adding circuit breaker configurations for cache timeouts and using more resilient cache options. We have also decreased max request timeouts within the billing service that caused impact and added more capacity to our cache to prevent traffic spikes from having the same impact.

Apr 23, 15:02 UTC
Monitoring - The degradation affecting Actions, Codespaces, Copilot and Packages has been mitigated. We are monitoring to ensure stability.

Apr 23, 15:02 UTC
Update - A mitigation was applied and services have recovered.  Actions is working through queued work before fully recovering.

Apr 23, 14:51 UTC
Update - Users are experiencing errors loading various web pages on github.com. Actions and Copilot Cloud Agent runs will be delayed.

Apr 23, 14:44 UTC
Update - Copilot is experiencing degraded performance. We are continuing to investigate.

Apr 23, 14:42 UTC
Update - Codespaces is experiencing degraded performance. We are continuing to investigate.

Apr 23, 14:41 UTC
Update - Packages is experiencing degraded performance. We are continuing to investigate.

Apr 23, 14:40 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
GitHubApr 23, 2026 at 05:00 UTC

Disruption with some GitHub services

Apr 23, 05:00 UTC
Resolved - Between 06:07 and 06:57 UTC, during a routine OS upgrade to our search infrastructure, a pre-existing corruption on one of our coordinator nodes in a single availability zone caused it to stop serving requests. As a result, search query requests routed to that availability zone were dropped, affecting approximately 142,000 requests over the course of the incident. The issue was resolved by provisioning additional coordinator nodes in the availability zone to restore capacity, after which the corrupted node was decommissioned.

View →
VercelresolvedApr 23, 2026 at 01:12 UTC

Delays Loading Logs

Apr 23, 01:12 UTC
Resolved - This incident has been resolved.

Apr 23, 00:51 UTC
Monitoring - A fix has been implemented for the issue causing delays in loading logs and we are seeing signs of recovery. We'll provide additional updates as-needed.

Apr 22, 23:08 UTC
Identified - We've identified an additional issue causing delays in loading logs. We'll provide additional updates as they become available.

Apr 22, 22:25 UTC
Monitoring - A fix has been implemented for the issue causing delays in loading logs and we are seeing signs of recovery. We'll provide additional updates as-needed.

Apr 22, 21:41 UTC
Update - We are continuing to implement a fix for the issue causing delays in loading logs. We will provide additional updates as they become available.

Apr 22, 19:31 UTC
Update - We are continuing to implement a fix for the issue causing delays in loading logs. We will provide additional updates as they become available.

Apr 22, 19:02 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 22, 18:25 UTC
Investigating - We're investigating an issue causing delays in loading logs. We will provide additional updates as they become available.

View →
VercelresolvedApr 22, 2026 at 23:00 UTC

Partial Observability Outage Affecting Compute Telemetry

Apr 22, 23:00 UTC
Resolved - We've identified an issue causing missing observability data for compute products between April 22, 2026 23:00 UTC and April 23, 2026 05:00 UTC affecting a subset of customers, resulting in missing or incomplete metrics in Observability dashboard during this period.

All underlying compute services - including Vercel Functions, Workflows, and Cron Jobs - continued to operate normally. Logs were not affected and remain fully available.

View →
GitHubresolvedApr 22, 2026 at 22:43 UTC

Disruption with some GitHub services

Apr 22, 22:43 UTC
Resolved - On April 22, 2026, between 09:00 UTC and 22:05 UTC, the Copilot coding agent and pull request comment event processing were degraded. During this period, approximately 0.5% of total pull request and issue comments mentioned @copilot (~23,000 invocations), explicitly requested work from the Copilot coding agent but were not acted upon.

Creating, viewing, and replying to pull request comments was unaffected, and other Copilot
functionality continued to operate normally. The impact was limited to @copilot mentions on pull request comments not triggering Copilot coding agent runs, and to some downstream systems not receiving new pull request comment events during the impact window.

The cause was a serialization error that prevented pull request comment events from being published to downstream consumers, including the Copilot coding agent. This was related to the same class of issue as incident #4295 on April 20, affecting a another event type.

We mitigated the incident by deploying a fix that restored event publishing, after which the Copilot coding agent and other downstream consumers resumed processing pull request comment events normally.

We are working to complete our audit of related event schemas, migrate remaining consumers to use
the updated identifier fields, and improve monitoring to detect drops in publishing on critical event topics, to reduce our time to detection and mitigation of issues like this one in the future.

Apr 22, 22:09 UTC
Update - We have identified the root cause of the disruption affecting Copilot Coding Agent and Issues. A fix is being deployed.

Apr 22, 20:37 UTC
Update - We have identified the root cause of the disruption affecting Copilot Coding Agent and Issues. Copilot @-mentions on pull requests are not being processed, and some issue-related functionality may be degraded. A fix has been developed and is being applied.

Apr 22, 20:02 UTC
Update - Copilot @-mentions on pull requests are currently not being processed by Copilot Cloud Agent. We have found the issue and are investigating remediations.

Apr 22, 19:55 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Apr 22, 19:53 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 22, 2026 at 19:18 UTC

Disruption with Copilot chat and Copilot Coding Agent

Apr 22, 19:18 UTC
Resolved - On April 22, 2026, between 15:16 UTC and 19:18 UTC, users experienced errors when interacting with Copilot Chat on github.com and Copilot Cloud Agent. During this time, users were unable to use Copilot Chat or Copilot Cloud Agent. Copilot Memory (in preview) was not available to Copilot agent sessions during this time. The issue was caused by an infrastructure configuration change that resulted in connectivity issues with our databases. The team identified the cause and restored connectivity to the database. Copilot Chat and Cloud Agent for github.com were restored by 18:16 UTC. Remaining regional deployments were restored incrementally, with full resolution at 19:18 UTC. We have taken steps to prevent similar infrastructure changes from causing these kinds of database operations in the future.

Apr 22, 18:05 UTC
Update - Copilot cloud agent and chat are mitigated for github.com.

Apr 22, 17:49 UTC
Update - We are now seeing recovery for Copilot cloud agent.

Apr 22, 17:40 UTC
Update - Mitigation is progressing for Copilot chat and cloud agent recovery.

Apr 22, 16:58 UTC
Update - Mitigation is progressing for Copilot chat and cloud agent.

Apr 22, 16:24 UTC
Update - We continue to work on mitigation for Copilot chat and cloud agent.

Apr 22, 15:43 UTC
Update - We are aware of users seeing errors interacting with Copilot chat on github.com and Copilot cloud agent. We have identified the cause and are investigating remediations.

Apr 22, 15:35 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 22, 2026 at 01:24 UTC

Disruption with projects service

Apr 22, 01:24 UTC
Resolved - On April 21, 2026, between 13:35 UTC and 01:24 UTC the following day the projects service was degraded. During this time period, projects may have been out of sync and users may have experienced delays in changes to projects and their items. Delays in reflected changes peaked at approximately 45 minutes. The delays were caused by serialization errors that failed events and triggered a flood of resyncs, overloading our event processing layers.

We mitigated the incident by speeding up processing time for incoming changes and otherwise waiting for all changes to be processed.

We are working to increase our capacity for processing updates to projects to reduce our time to mitigation of issues like this one in the future.

Apr 22, 00:00 UTC
Update - The issue remains mitigated. Users may still experience small delays in changes to projects while we process the backlog of events. We expect a full recovery in approximately two hours.

Apr 21, 22:49 UTC
Update - The issue remains mitigated. Users may still experience delays in changes to projects while we process the backlog of events. We expect a full recovery in approximately three hours.

Apr 21, 21:15 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 21, 19:41 UTC
Update - Recovery from the delays affecting GitHub Projects continues to progress. We have deployed additional mitigations that are accelerating processing of the backlog. Users may still experience delays where changes to projects are not reflected immediately. We expect full recovery within approximately six hours.

Apr 21, 17:21 UTC
Update - The queues are continuing to decrease and we are working to accelerate the rate of processing through the queues.

Apr 21, 16:45 UTC
Update - The mitigation is deployed and we are seeing recorvery in the queues and will provide an update as to when full recovery will be realized.

Apr 21, 16:18 UTC
Update - We are deploying a fix to relieve the queue of delayed data. Some users may still experience delays with GitHub Projects where changes are not reflected immediately as remaining backlogs are processed.

Apr 21, 15:42 UTC
Update - We continue to investigate delays with GitHub Projects where changes may not be reflected immediately. Our team has identified the cause and applied mitigations to address the issue. We are seeing initial signs of recovery, though some delays may persist as the system works through a backlog of pending updates.

Apr 21, 15:07 UTC
Update - We are investigating reports of delays with GitHub Projects. Users may notice that changes made to projects are not reflected immediately. Our team has identified the source of the delays and is actively working to resolve the issue.

Apr 21, 15:03 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 21, 2026 at 14:41 UTC

Elevated errors on Vercel Dashboard

Apr 21, 14:41 UTC
Resolved - This incident has been resolved.

Apr 21, 14:30 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Apr 21, 14:21 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 21, 13:51 UTC
Investigating - We are currently investigating this issue.

View →
GitHubresolvedApr 21, 2026 at 05:04 UTC

Partial degradation for code scanning default setup and for code quality

Apr 21, 05:04 UTC
Resolved - On April 20, 2026 between 10:28 UTC and 15:04 UTC GitHub experienced degraded service for code scanning default setup, code quality, and project boards. Repair of affected project boards additionally lasted until April 21, 05:04 UTC

During this time, code scanning default setup and code quality analyses were not triggered on newly opened pull requests. Additionally, newly created issues were not appearing on project boards.

The cause was a serialization error that prevented proper triggering of code scanning, code quality analyses, and project board updates.

We mitigated the issue by deploying a fix, restoring event publishing for code scanning and code quality. For project boards, an additional code change was deployed to update event consumers, followed by a reindex of affected project items.

We are working to prevent recurrence by strengthening our schema validations and improving monitoring for drops in publishing on critical hydro topics.

Apr 21, 04:18 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 21, 03:10 UTC
Update - The issue remains mitigated. Issues that were linked to projects during the incident may take approximately three more hours to render correctly while we complete a re-index.

Apr 20, 21:36 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 20, 18:21 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 20, 18:20 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 20, 18:20 UTC
Update - The issue has been mitigated. Newly created issues linked to projects should now function as expected. Issues that were linked to projects during the incident may take approximately five hours to render correctly while we complete a re-index.

Apr 20, 18:08 UTC
Update - A deployment to fix this issue of new issues not showing up in projects is underway.

Apr 20, 17:32 UTC
Update - We continue to work on mitigation regarding new issues not showing on project boards.

Apr 20, 16:48 UTC
Update - We continue to work on mitigation regarding new issues not showing on project boards.

Apr 20, 16:16 UTC
Update - Code scanning default setup and Code Quality triggers are back up and running. PRs not processed before or during this incident will require a new push to trigger code scanning or code quality analysis.

We are seeing problems with new issues not showing on project boards and are working on mitigation.

Apr 20, 15:20 UTC
Update - We are continuing to work on a mitigation to unblock code scanning default setup and code quality features on pull requests.

Apr 20, 14:38 UTC
Update - We are currently deploying mitigations that should unblock code scanning default setup and code quality features on pull requests.

Apr 20, 13:57 UTC
Update - We are actively working to mitigate an issue affecting code scanning default setup and code quality features on pull requests. Users may experience pull request code scanning and code quality analyses not being triggered on new pull requests. Our engineering team has identified the root cause and working on mitigating the issue.

Apr 20, 13:28 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 20, 2026 at 22:18 UTC

Observability Alerts degraded

Apr 20, 22:18 UTC
Resolved - This incident has been resolved.

Apr 20, 22:08 UTC
Monitoring - We have implemented a fix targeted at usage anomalies and are now observing recovery across all anomaly alerts. We will share any additional updates as-needed.

Apr 20, 21:34 UTC
Update - We are continuing to work on a fix for usage anomalies. We'll provide additional updates as they become available.

Apr 20, 18:50 UTC
Update - We are continuing to work on a fix for usage anomalies. We'll provide additional updates as they become available.

Apr 20, 16:51 UTC
Update - We have resolved the issue for error anomalies, and are continuing to work on a fix for usage anomalies.

Apr 20, 14:10 UTC
Update - We are continuing to work on a fix for this issue.

Apr 20, 12:21 UTC
Update - We are continuing to work on a fix for this issue.

Apr 20, 10:40 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 20, 09:38 UTC
Investigating - We are currently investigating Observability Alerts not detecting new events. We will share more updates as soon as we have more information.

View →
VercelresolvedApr 20, 2026 at 20:56 UTC

Elevated errors on Vercel Dashboard and API endpoints

Apr 20, 20:56 UTC
Resolved - This incident has been resolved.

Apr 20, 20:33 UTC
Monitoring - A fix has been implemented and we are observing recovery in error rates across the Vercel Dashboard and API endpoints. We will continue to monitor and share any additional updates as they become available.

Apr 20, 19:40 UTC
Identified - We have identified an issue causing elevated error rates on the Vercel Dashboard and API endpoints. Some users may experience an issue creating new deployments, but existing deployments are not impacted. We will share more updates as they become available.

View →
VercelresolvedApr 20, 2026 at 10:36 UTC

Unaccessible runtime logs in the dashboard

Apr 20, 10:36 UTC
Resolved - This incident has been resolved.

Apr 20, 10:26 UTC
Monitoring - A fix has been implemented and we are monitoring the result. Please refresh your browser page if you are not seeing runtime logs.

Apr 20, 10:10 UTC
Investigating - We are currently investigating runtime logs not loading properly in the dashboard. We will share more updates as soon as we have more information.

View →
VercelresolvedApr 19, 2026 at 21:39 UTC

Delays Loading Observability, Analytics and Speed Insights

Apr 19, 21:39 UTC
Resolved - This incident has been resolved.

Apr 19, 21:22 UTC
Monitoring - A fix has been implemented and we are seeing recovery. No action is required - data may continue to be partially delayed as it is backfilled.

Apr 19, 21:12 UTC
Identified - The issue has been identified and a fix is being implemented.

Apr 19, 20:43 UTC
Investigating - We are currently experiencing delays in loading Observability, Analytics and Speed Insights data on the dashboard. This may also affect User Activity Events and Firewall UI. Vercel Alerts may also be delayed. We will provide additional updates as they become available.

View →
GitHubresolvedApr 17, 2026 at 15:18 UTC

Disruption with some GitHub services

Apr 17, 15:18 UTC
Resolved - On April 17, 2026, between 14:46 UTC and 15:12 UTC, users experienced a degraded web experience on GitHub.com. During this time, approximately 1.5% of web requests resulted in errors, with some users encountering slow page loads or failed requests. The issue was caused by capacity saturation of a caching component in one of our data center regions. We mitigated the issue by redirecting traffic to an unaffected region and rolling back a recent deployment. The incident was fully resolved at 15:18 UTC. We are taking steps to provide appropriate capacity for this caching path to prevent recurrence.

Apr 17, 15:18 UTC
Monitoring - The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

Apr 17, 15:08 UTC
Update - We have isolated a problematic component in our infrastructure and are working to mitigate. We will continue to post updates as we work toward resolution.

Apr 17, 14:57 UTC
Update - We are experiencing an issue that impacts approximately 10% of traffic to the web, resulting in slow and failed calls. We are investigating and will continue to post updates as we work toward mitigation.

Apr 17, 14:56 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Apr 17, 14:56 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 16, 2026 at 18:28 UTC

Incident with Codespaces

Apr 16, 18:28 UTC
Resolved - On April 16, 2026 between 09:30 UTC and 17:15 UTC, users experienced failures when attempting to connect to GitHub Codespaces via the VS Code editor. During this time, approximately 40% of codespace start operations failed. Users connecting via SSH were not impacted.

The issue was caused by a failure in an upstream download service that prevented the VS Code Server from being retrieved during codespace startup. The impact was mitigated by implementing a workaround to use an alternative download path when the primary endpoint is degraded.

We are working with the upstream dependency to address the root cause of the download service failure, and we are improving our fallback mechanisms to reduce the impact of similar upstream failures in the future.

Apr 16, 18:22 UTC
Monitoring - The degradation affecting Codespaces has been mitigated. We are monitoring to ensure stability.

Apr 16, 16:37 UTC
Update - Our provider is implementing a mitigation and we are seeing signs of recovery.

Apr 16, 15:49 UTC
Update - We found an issue that impacts 70% of Codespaces. We are engaged with the provider and working towards mitigation.

Apr 16, 15:41 UTC
Update - Codespaces is experiencing degraded availability. We are continuing to investigate.

Apr 16, 15:08 UTC
Update - We are experiencing degraded performance in Codespaces related to creating a new Codespace or starting an existing Codespace from the VS Code editor. SSH connections to Codespaces are not impacted. We are working toward mitigation and will continue to keep you updated on progress.

Apr 16, 15:06 UTC
Investigating - We are investigating reports of degraded performance for Codespaces

View →
InngestApr 15, 2026 at 23:21 UTC

Delays in function run scheduling

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
  • Observability (Operational)
View →
ResendApr 15, 2026 at 17:05 UTC

Increased latency for broadcast sending and domain verification

Status: Resolved

The pending domain verifications and broadcasts are catching-up quickly. The incident should fully resolve soon.

Affected components
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • General API (Operational)
View →
GitHubresolvedApr 14, 2026 at 06:08 UTC

Disruption with some GitHub services

Apr 14, 06:08 UTC
Resolved - On April 14, between 00:58 UTC and 06:08 UTC, GitHub Enterprise Cloud customers experienced 500 errors when attempting to access Copilot Insights pages which was caused by an authentication failure in our metrics pipeline. We fully mitigated the issue and validated the fix in production. Approximately 709 users were impacted. The total impact duration was approximately 5 hours and 10 minutes.

Our investigation determined the incident was caused by a change in a tenant credential which caused authentication errors to retrieve the required data needed on our Copilot Insights pages.

We understand this disruption impacted customers' ability to access the Copilot Insights page. To prevent similar issues and reduce resolution time in the future, we are investing in improved diagnostics tooling to quickly identify the root cause of failures, enhanced monitoring, and alerting to detect issues at a more granular level.

GitHub is a critical infrastructure for your work, your teams, and your businesses. We are focused on these remediations and continued reliability improvements for Copilot Insights and related metrics experiences.

Apr 14, 06:07 UTC
Update - This incident has been resolved. We will continue to monitor to ensure stability. Thank you for your patience and understanding as we addressed this issue.

Apr 14, 06:07 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 14, 04:40 UTC
Update - We identified an issue that impacts the Copilot Dashboard on the Insights tab and are working on mitigation. We will continue to keep you updated on progress.

Apr 14, 03:47 UTC
Update - The team continues to investigate issues accessing with Copilot Dashboard on the Insights tab.  We will continue providing updates on the progress towards mitigation.

Apr 14, 02:40 UTC
Update - The Copilot Dashboard on the Insights tab is not accessible and we are continuing to investigate.

Apr 14, 02:37 UTC
Update - Degradation of Service - Insights Page

Apr 14, 01:57 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 13, 2026 at 20:35 UTC

Incident with Pages

Apr 13, 20:35 UTC
Resolved - On Sunday April 13th, 2026, between 18:53 UTC and 20:30 UTC, the GitHub Pages service experienced elevated error rates. On average, the error rate was 10.58% and peaked at 12.77% of requests to the service, resulting in approximately 17.5 million failed requests returning HTTP 500 errors. This was due to an automated DNS management tool (octodns) erroneously deleting a DNS record for a Pages backend storage host after its upstream data source intermittently failed to return the record, causing the tool to treat it as stale and remove it.

We mitigated the incident by re-creating the deleted DNS record. To prevent future incidents, we are implementing availability-zone-tolerant routing in the Pages frontend so that an unresolvable backend host triggers failover to healthy hosts rather than returning errors, adding safeguards to prevent automated deletion of DNS records owned by other systems, and improving logging and alerting for DNS resolution failures in the Pages serving path.

Apr 13, 20:32 UTC
Update - We have mitigated the issue with Pages.

Apr 13, 20:30 UTC
Monitoring - The degradation affecting Pages has been mitigated. We are monitoring to ensure stability.

Apr 13, 19:57 UTC
Update - We are investigating reports of issues with Pages. We will continue to keep users updated on progress towards mitigation.

Apr 13, 19:56 UTC
Investigating - We are investigating reports of degraded availability for Pages

View →
GitHubresolvedApr 13, 2026 at 17:40 UTC

Disruption with some GitHub services

Apr 13, 17:40 UTC
Resolved - On April 13, 2026, between 14:41 UTC and 17:29 UTC, the Copilot service experienced degraded performance. All Copilot users were impacted by increased latency, and approximately 20% experienced request failures when interacting with Copilot Cloud Agent (CCA). On average, request latency increased to approximately 950ms. The GitHub User Dashboard also displayed intermittent errors loading Copilot quota information. CCA and the User Dashboard were impacted for approximately 2 hours and 56 minutes.

This was due to an infrastructure change that reduced the available compute capacity for a backend service responsible for Copilot rate limiting and quota management. The reduced capacity caused resource exhaustion under normal traffic load, leading to cascading failures in downstream request processing.

We mitigated the incident by increasing compute resources allocated to the affected service and scaling out the number of service instances to distribute load more effectively.

We are working to improve proactive capacity monitoring to detect resource degradation before it impacts users, reviewing retry and timeout configurations across dependent services to reduce amplification during degraded states, and evaluating connection management strategies to improve resilience under constrained resources.

Apr 13, 16:59 UTC
Update - We have identified the root cause and are rolling out a fix for Copilot. The services should now be in recovery, with expected full recovery in 5 to 10 minutes.

Apr 13, 16:41 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 10, 2026 at 13:28 UTC

Problems with third-party Claude and Codex Agent sessions not being listed in the agents tab dashboard

Apr 10, 13:28 UTC
Resolved - On April 9, 2026, between 22:59 UTC and April 10, 2026, 13:24 UTC, the Copilot Mission Control service was degraded and did not display Claude and Codex Cloud Agent sessions in the agents tab dashboard. Customers were unable to see, list, or manage their third party agent sessions during this period. The underlying agent sessions continued to function normally. This was a visibility and management issue only, and no HTTP errors were generated. The API returned successful responses with incomplete results, with an average error rate of 0% and a maximum error rate of 0%. This was due to a code change that introduced a filter which inadvertently excluded third party agent sessions.

We mitigated the incident by reverting the problematic code change and deploying the fix to production.

We are working to add automated monitoring for dashboard content visibility and improve integration test coverage for third party agent session listing to reduce our time to detection and mitigation of issues like this one in the future.

Apr 10, 13:08 UTC
Update - We are investigating third party Claude and Codex Cloud Agent sessions not being listed in the agents tab dashboard.

Apr 10, 13:07 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendApr 10, 2026 at 12:20 UTC

Email sending error

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Batch Emails (Operational)
  • General API (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
View →
GitHubresolvedApr 9, 2026 at 20:36 UTC

Disruption with some GitHub services

Apr 9, 20:36 UTC
Resolved - On April 9, 2026, between 16:05 UTC and 20:36 UTC, the Copilot cloud agent service was degraded, causing new agent sessions to be delayed or fail to start. Users who attempted to start Copilot cloud agent sessions during this period experienced jobs getting stuck in the queue, with wait times peaking at 54 minutes compared to the normal 15–40 seconds. On average, approximately 84% of requests to start agent sessions failed, peaking at 97.5% during the worst period.

This was due to an internal service exceeding API rate limits, compounded by a caching bug that persisted the rate-limited state beyond the actual rate limit window, causing recurring outage waves rather than a single recovery.

We mitigated the incident by deploying a configuration change to bypass the affected cache and shifting API traffic to an alternative authentication path that reduced rate limit exposure. We have since added automated monitoring and alerting for this failure mode, deployed per-endpoint rate limit controls, and added caching for high-traffic API calls to reduce overall load. We are also working on longer-term improvements to rate limit isolation and traffic management to prevent similar issues in the future.

This incident shared the same underlying root causes with an incident declared in the time frame https://www.githubstatus.com/incidents/zn1t56bfxdzg

Apr 9, 19:52 UTC
Update - We continue to investigate periodic delays in Copilot Cloud Agent job processing

Apr 9, 18:57 UTC
Update - We are continuing to investigate Copilot Cloud Agent job delays

Apr 9, 17:48 UTC
Update - Copilot Cloud Agent jobs are being processed and we are monitoring recovery

Apr 9, 16:57 UTC
Update - We are investigating delays processing Copilot Cloud Agent jobs

Apr 9, 16:20 UTC
Update - We are experiencing issues where jobs are being delayed to start for copilot coding agent

Apr 9, 16:20 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 9, 2026 at 10:15 UTC

Disruption with some GitHub services

Apr 9, 10:15 UTC
Resolved - On April 9, 2026, between 09:05 UTC and 19:05 UTC, the Copilot coding agent service was degraded and users experienced significant delays starting new agent sessions. Approximately 84% of new agent session requests were delayed across four separate outage waves, with queue wait times peaking at 54 minutes compared to a normal baseline of 15–40 seconds. On average, the error rate was 83.9% and peaked at 97.5% of requests to the service. Approximately 22,700 workflow creations were delayed or failed during the incident.

This was due to a bug in our rate limiting logic that incorrectly applied a rate limit globally across all users, rather than scoping it to the individual installation that triggered the limit. A contributing factor was a surge in API traffic from a client update that increased requests to an internal endpoint by 3–4x, which accelerated rate limit exhaustion.

We mitigated the incident by disabling the faulty rate limit caching mechanism via feature flag and updating our service to use per-installation credentials for API calls, ensuring rate limits are correctly scoped to individual installations.

We have since added automated monitoring and alerting to detect this failure mode proactively, deployed fixes to reduce unnecessary API traffic through caching improvements, and are continuing work to further isolate rate limit scoping across client types to prevent similar issues in the future.

This incident shared the same underlying root causes with an incident declared in the time frame https://www.githubstatus.com/incidents/2rqwxl8y7m0j

Apr 9, 10:15 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 9, 09:57 UTC
Update - We are investigating an issue affecting GitHub Copilot coding agent. Users may experience significant delays when starting new agent sessions, with jobs remaining queued longer than expected. Our team has identified increased load as a contributing factor and is actively working to restore normal performance.

Apr 9, 09:50 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 9, 2026 at 04:57 UTC

Disruption with GitHub notifications

Apr 9, 04:57 UTC
Resolved - On April 9, 2026, between 03:22 UTC and 04:49 UTC, GitHub Notifications experienced degraded availability. During this time, approximately 45% of requests to the notifications service returned errors, with a peak error rate of approximately 54%, preventing affected users from successfully viewing or interacting with their notifications service. The issue was identified and resolved, restoring the service to full availability.

We are working to improve our metrics to reduce time to detection and mitigation for similar issues in the future.

Apr 9, 04:57 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 9, 04:42 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 7, 2026 at 18:56 UTC

Elevated queue times for deployments on the Hobby plan

Apr 7, 18:56 UTC
Resolved - This incident has been resolved.

Apr 7, 18:27 UTC
Update - A fix has been rolled out, and we are monitoring the results. Queued deployments are being processed and new deployments are building normally now.

Apr 7, 18:06 UTC
Identified - We have identified an issue causing elevated queue times for deployments on the Hobby plan. We are investigating a fix and will share more updates as soon as we have more information.

View →
VercelresolvedApr 7, 2026 at 08:37 UTC

Builds logs stuck in loading state in iad1

Apr 7, 08:37 UTC
Resolved - The incident has been resolved. If you don't see build logs for a specific deployment, please redeploy.

Apr 7, 08:29 UTC
Monitoring - We identified the cause of build logs stuck in the loading state affecting iad1 and are working around the issue for new builds. Please redeploy to your build to view build logs.

Apr 7, 07:19 UTC
Investigating - We are investigating reports of build logs stuck in loading state.

View →
ResendApr 3, 2026 at 06:25 UTC

Delay on sending emails

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Batch Emails (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
View →
VercelresolvedApr 2, 2026 at 23:09 UTC

Elevated INTERNAL_UNEXPECTED_ERROR error rates for some deployments

Apr 2, 23:09 UTC
Resolved - This incident has been resolved.

Apr 2, 20:02 UTC
Monitoring - The fix has been rolled out and we are monitoring the results. If you are still seeing issues, we recommend redeploying or performing an instant rollback to fix the issue.

Apr 2, 19:26 UTC
Identified - We have identified the issue and are rolling out a fix. Deployments created between 16:10 to 19:12 UTC may be affected. If you are still seeing issues, we recommend redeploying or performing an instant rollback to fix the issue.

Apr 2, 18:39 UTC
Investigating - We are investigating reports of elevated INTERNAL_UNEXPECTED_ERROR responses affecting some customer deployments. We will provide updates as we learn more.

View →
GitHubresolvedApr 2, 2026 at 21:48 UTC

Disruption with some GitHub services

Apr 2, 21:48 UTC
Resolved - Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected.

Once the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards.

Apr 2, 21:48 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 2, 20:35 UTC
Update - Although we are observing recovery once again, we expect continued periods of degradation.

Work that is queued during times of degradation does eventually get processed.

We continue to investigate and find a mitigation, and will update again within 2 hours.

Apr 2, 19:28 UTC
Update - This issue has recurred. Customers will once again experience false job starts when assigning tasks to Copilot Cloud Agent.

We are still investigating and trying to understand the pattern of degradation.

Apr 2, 18:25 UTC
Update - We are once again seeing recovery with Copilot Cloud Agent job starts.

We are keeping this open while we verify this won't recur.

Apr 2, 17:59 UTC
Update - When assigning tasks to Copilot Cloud Agent, the task will appear to be working, but may not actually be running.

We are investigating.

Apr 2, 17:49 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 2, 2026 at 18:53 UTC

Elevated errors (SIGSEGV) for Vercel Functions running Node.js 20 in cle1 and dub1 regions

Apr 2, 18:53 UTC
Resolved - This incident has been resolved.

Apr 2, 18:09 UTC
Monitoring - The root cause has been identified, a fix has been rolled out, and we are actively monitoring. Newly deployed Vercel Functions on Node.js 20 in cle1 and dub1 no longer terminate with SIGSEGV. We recommend redeploying affected Vercel Functions if you are still seeing issues.

Apr 2, 16:13 UTC
Update - We are continuing to investigate increased rates of Vercel Functions running on Node.js 20 terminating with SIGSEGV in Vercel's cle1 and dub1 regions. Other regions are unaffected. We recommend redeploying affected Vercel Functions or upgrading them to either Node js 22 or 24.

Apr 2, 14:10 UTC
Investigating - We are currently investigating increased rates of Vercel Functions running on Node.js 20 terminating with SIGSEGV. We recommend redeploying affected Vercel Functions or upgrading them to either Node.js 22 or 24.

View →
InngestApr 2, 2026 at 17:49 UTC

Dashboard down

Status: Resolved

The incident is now resolved and the system is full operational. Related to the Vercel incident, we updated our dashboard to Node 22.x to solve the issue. We continue to monitor Vercel's incident and react accordingly. https://www.vercel-status.com/incidents/5r9bp5y8rql2

Affected components
  • Inngest Dashboard (Operational)
View →
GitHubresolvedApr 2, 2026 at 16:30 UTC

Copilot Coding Agent failing to start some jobs

Apr 2, 16:30 UTC
Resolved - Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected.

Once the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards.

This was the same incident declared in https://www.githubstatus.com/incidents/d96l71t3h63k

Apr 2, 16:28 UTC
Update - When assigning tasks to Copilot Cloud Agent, the task will appear to be working, but may not actually be running.

We are investigating.

Apr 2, 16:18 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
ResendApr 2, 2026 at 15:05 UTC

Email Sending Delay

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • Webhooks (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Website (Operational)
View →
GitHubresolvedApr 1, 2026 at 23:45 UTC

Disruption with GitHub's code search

Apr 1, 23:45 UTC
Resolved - On April 1st, 2026 between 14:40 and 17:00 UTC the GitHub code search service had an outage which resulted in users being unable to perform searches.

The issue was initially caused by an upgrade to the code search Kafka cluster ZooKeeper instances which caused a loss of quorum. This resulted in application-level data inconsistencies which required the index to be reset to a point in time before the loss of quorum occurred. Meanwhile, an accidental deploy resulted in query services losing their shard-to-host mappings, which are typically propagated by Kafka.

We remediated the problem by performing rolling restarts in the Kafka cluster, allowing quorum to be reestablished. From there we were able to reset our index to a point in time before the inconsistencies occurred.

The team is working on ways to improve our time to respond and mitigate issues relating to Kafka in the future.

Apr 1, 23:45 UTC
Update - Code search has recovered and is serving production traffic.

Apr 1, 22:00 UTC
Update - We have stabilized Code Search infrastructure, and are in the final stages of validation before slowly reintroducing production traffic.

Apr 1, 19:37 UTC
Update - We are still working on recovering back to a serviceable state and expect to have a more substantial update within another two hours.

Apr 1, 17:48 UTC
Update - We are observing some recovery for Code Search queries, but customers should be aware that the data being served may be stale, especially for changes that took place after 07:00 UTC today (1 April 2026). We are still working on recovering our ingestion pipeline, and synchronizing the indexed data.

We will update again within 2 hours.

Apr 1, 16:00 UTC
Update - We identified an issue in our ingestion pipeline that degraded the freshness of Code Search results. While fixing the issue with the ingestion pipeline, a deployment caused a loss of dynamic configuration which is causing most requests for Code Search results to fail. We are working to restore the service and to re-ingest the misaligned data.

Apr 1, 15:02 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedApr 1, 2026 at 22:17 UTC

Elevated errors creating deployments with integrations and increased delays in delivering webhooks

Apr 1, 22:17 UTC
Resolved - This incident has been resolved.

Apr 1, 21:57 UTC
Monitoring - A fix has been implemented and we are monitoring the results. Webhook events are being delivered normally, and deployments with integrations are building successfully.

Apr 1, 20:48 UTC
Identified - We have identified an issue causing elevated latency and errors with delivering webhooks and creating deployments that use integrations. We are working on a fix for this issue and will provide further updates as soon as they become available.

View →
GitHubresolvedApr 1, 2026 at 16:10 UTC

GitHub audit logs are unavailable

Apr 1, 16:10 UTC
Resolved - On April 1, 2026, between 15:34 UTC and 16:02 UTC, our audit log service lost connectivity to its backing data store due to a failed credential rotation. During this 28-minute window, audit log history was unavailable via both the API and web UI. This resulted in 5xx errors for 4,297 API actors and 127 github.com users. Additionally, events created during this window were delayed by up to 29 minutes in github.com and event streaming. No audit log events were lost; all audit log events were ultimately written and streamed successfully. Customers using GitHub Enterprise Cloud with data residency were not impacted by this incident.

We were alerted to the infrastructure failure at 15:40 UTC — six minutes after onset — and resolved the issue by recycling the affected environment, restoring full service by 16:02 UTC. We are conducting a thorough review of our credential rotation process to strengthen its resiliency and prevent recurrence. In parallel, we are strengthening our monitoring capabilities to ensure faster detection and earlier visibility into similar issues going forward.

Apr 1, 16:07 UTC
Update - A routine credential rotation has failed for our our audit logs service; we have re-deployed our service and are waiting for recovery.

Apr 1, 16:06 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
GitHubresolvedApr 1, 2026 at 12:41 UTC

Incident with Copilot

Apr 1, 12:41 UTC
Resolved - On April 1, 2026, between 07:29 and 12:41 UTC, some customers experienced elevated 5xx errors and increased latency when using GitHub Copilot features that rely on `/agents/sessions` endpoints (including creating or viewing agent sessions). The issue was caused by resource exhaustion in one of the Copilot backend services handling these requests, in turn, causing timeouts and failed requests. We mitigated the incident by increasing the service’s available compute resources and tuning its runtime concurrency settings. Service health returned to normal and the incident was fully resolved by 12:41 UTC.

Apr 1, 12:10 UTC
Update - The success rate and latency for creating and viewing agent sessions has stabilized at baseline levels, we are continuing to monitor recovery

Apr 1, 12:02 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 1, 11:37 UTC
Update - The success rate for creating and viewing agent sessions has stabilized, and we're continuing to monitor latency, which is trending toward baseline levels.

Apr 1, 11:24 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Apr 1, 10:56 UTC
Monitoring - The degradation affecting Copilot has been mitigated. We are monitoring to ensure stability.

Apr 1, 10:31 UTC
Update - Users may see increased latency and intermittent errors when viewing or creating agent sessions. We are working on mitigations to return to baseline performance and success rate.

Apr 1, 10:00 UTC
Update - We are investigating reports of issues with service(s): Copilot Dotcom Agents. We will continue to keep users updated on progress towards mitigation.

Apr 1, 09:58 UTC
Investigating - We are investigating reports of degraded performance for Copilot

View →
GitHubMar 31, 2026 at 21:23 UTC

Incident with Pull Requests: High percentage of 500s

Mar 31, 21:23 UTC
Resolved - On Monday March 31st, 2026, between 13:53 UTC and 21:23 UTC the Pull Requests service experienced elevated latency and failures. On average, the error rate was 0.15% and peaked at 0.28% of requests to the service. This was due to a change in garbage collection (GC) settings for a Go-based internal service that provides access to Git repository data. The changes caused more frequent GC activity and elevated CPU consumption on a subset of storage nodes, increasing latency and failure rates for some internal API operations.

We mitigated the incident by reverting the GC changes. To prevent future incidents and improve time to detection and mitigation, we are instrumenting additional metrics and alerting for GC-related behavior, improving our visibility into other signals that could cause degraded impact of this type, and updating our best practices and standards for garbage collection in Go-based services.

Mar 31, 21:16 UTC
Monitoring - The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

Mar 31, 21:12 UTC
Update - We continue to see a small subset of repositories experiencing timeouts and elevated latency in Pull Requests, affecting under 1% of requests.

Mar 31, 19:28 UTC
Update - Error rates remain elevated across multiple pull request endpoints. We are pursuing multiple potential mitigations.

Mar 31, 18:42 UTC
Update - We continue to experience elevated error rates affecting Pull Requests. An earlier fix resolved one component of the issue, but some users may still encounter intermittent timeouts when viewing or interacting with pull requests. Our teams are actively investigating the remaining causes.

Mar 31, 17:16 UTC
Update - We identified an issue causing increased errors when accessing Pull Requests. The mitigation is being applied across our infrastructure and we will continue to provide updates as the mitigation rolls out.

Mar 31, 16:35 UTC
Update - We are seeing recovery in latency and timeouts of requests related to pull requests, even though 500s are still elevated. While we are continuing to investigate, we are applying a mitigation and expect further recovery after it is applied.

Mar 31, 16:15 UTC
Update - We are continuing to investigate increased 500 errors affecting GitHub services. You may experience intermittent failures when using Pull Requests and other features. We are actively working to identify and resolve the underlying cause.

Mar 31, 15:39 UTC
Update - We are investigating increased 500 errors affecting GitHub services. You may experience intermittent failures when using Pull Requests and other features. We are actively working to identify and resolve the underlying cause.

Mar 31, 15:06 UTC
Update - We are seeing a higher than average number of 500s due to timeouts across GitHub services. We have a potential mitigation in flight and are continuing to investigate.

Mar 31, 15:05 UTC
Investigating - We are investigating reports of degraded performance for Pull Requests

View →
VercelresolvedMar 31, 2026 at 21:23 UTC

Elevated Errors Creating Deployments

Mar 31, 21:23 UTC
Resolved - This incident has been resolved.

Mar 31, 21:17 UTC
Monitoring - A fix has been applied for an issue where deployments were erroring with "invalid request" and we are monitoring the results. We recommend redeploying any failed deployments.

Mar 31, 21:11 UTC
Investigating - We are investigating reports of elevated errors when creating deployments. Existing deployments and live traffic are unaffected. We will provide updates as they become available.

View →
GitHubMar 31, 2026 at 15:10 UTC

Issues with metered billing report generation

Mar 31, 15:10 UTC
Resolved - On March 31, 2026, between 06:15 UTC and 15:30 UTC, the GitHub billing usage reports feature was degraded due to reduced server capacity. Customers requesting billing usage reports and loading the top usage by organization and repository on the billing overview and usage pages were impacted. The average error rate for usage report requests was 15%, peaking at 98% over an eight-minute window. For the billing pages, an average of 56% of requests failed to load the top usage cards. The root cause was an increase in billing usage report requests with large datasets, which exhausted the capacity of the nodes responsible for reporting data. There was no impact on billing charges.

We mitigated the incident by adjusting our auto-scaling thresholds to better meet our capacity needs. We are working to improve our metrics to reduce time to detection and mitigation for similar issues in the future.

Mar 31, 15:01 UTC
Monitoring - The degradation has been mitigated. We are monitoring to ensure stability.

Mar 31, 14:59 UTC
Update - We have applied mitigations to a data store related to billing reports, and are seeing partial recovery to billing report generation. We continue to monitor for full recovery.

Mar 31, 14:56 UTC
Update - We are seeing a high number of 500s due to timeouts across GitHub services. We are redeploying some of our core services and we expect that this allow us to recover.

Mar 31, 14:39 UTC
Update - We're continuing to see high failure rates on billing report generation, and are working on mitigations for a data store related to billing reports.

Mar 31, 13:56 UTC
Update - We're seeing issues related to metered billing reports, intermittently affecting metered usage graphs and reports on the billing page. We have identified an issue with a data store, and are working on mitigations.

Mar 31, 13:47 UTC
Investigating - We are investigating reports of impacted performance for some GitHub services.

View →
VercelresolvedMar 31, 2026 at 13:29 UTC

Degraded Support Case Submission

Mar 31, 13:29 UTC
Resolved - This incident has been resolved.

Mar 31, 13:16 UTC
Update - Support case submission has been confirmed to be functioning normally. We will share updates as they become available.

Mar 31, 13:16 UTC
Monitoring - A fix has been implemented and we are monitoring the results. We will share updates as they become available.

Mar 31, 13:10 UTC
Investigating - We are currently investigating reports of elevated errors when submitting support cases. Some customers may experience timeouts when creating support tickets. We will share updates as they become available.

View →
InngestMar 31, 2026 at 11:59 UTC

Increased failures with step.fetch, step.ai.infer

Status: Resolved

The incident is now resolved and the system is full operational. During this incident step.fetch and step.ai.infer were failing due to a bug causing empty request bodies to be returned. The root cause was determined, the system was rolled back and a fix will be rolled out today.

Affected components
  • API (REST and GraphQL) (Operational)
View →
GitHubMar 30, 2026 at 13:25 UTC

Elevated delays in Actions workflow runs and Pull Request status updates

Mar 30, 13:25 UTC
Resolved - On March 30, 2026, between 10:11 UTC and 13:25 UTC, GitHub Actions experienced degraded performance. During this time, approximately 2.65% of workflow jobs triggered by pull request events experienced start delays exceeding 5 minutes. The issue was caused by replication lag on an internal database cluster used by Actions, which triggered write throttling in our database protection layer and slowed job queue processing.

The replication lag originated from planned maintenance to scale the internal database. Newly added database hosts triggered guardrails in the throttling layer, restricting write throughput. The incident was mitigated by excluding the new hosts from replication delay calculations.

To prevent recurrence, we have updated our maintenance procedures to ensure new hosts are excluded from throttling assessments during scaling operations. Additionally, we are investing in automation to streamline this type of maintenance activity.

Mar 30, 13:25 UTC
Update - The degradation has been mitigated. We are monitoring to ensure stability.

Mar 30, 13:20 UTC
Monitoring - The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability.

Mar 30, 13:02 UTC
Investigating - We are investigating reports of degraded performance for Actions and Pull Requests

View →
ResendMar 28, 2026 at 04:02 UTC

Email sending delay

Status: Resolved

We have resolved the issue causing increased latency in email sending. Emails are being delivered normally.

Affected components
  • Single Email (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
View →
GitHubMar 27, 2026 at 05:00 UTC

Incident with Copilot

Mar 27, 05:00 UTC
Resolved - On March 27, 2026, from 02:30 to 04:56 UTC, a misconfiguration in our rate limiting system caused users on Copilot Free, Student, Pro, and Pro+ plans to experience unexpected rate limit errors. The configuration that was incorrectly applied was intended solely for internal staff testing of rate-limiting experiences. Copilot Business and Copilot Enterprise accounts were not affected.

During this period, affected users received error messages instructing them to retry after a certain time. Approximately 32% of active Free users, 35% of active Student users, 46% of active Pro users, and 66% of active Pro+ users were affected.

After identifying the root cause, we reverted the change and restored the expected rate limits. We are reviewing our deployment and validation processes to help ensure configurations used for internal testing cannot be inadvertently applied to production environments.

View →
InngestMar 26, 2026 at 23:45 UTC

Function run scheduling delays

Status: Resolved

The incident is now resolved and the system is full operational. This was related to an issue caused by the part of the system powering the debounce feature. The internal event backlog is fully caught up and the two mitigations deployed have addressed the issue. The team is preparing a post-mortem to ensure this issue does not reoccur.

Affected components
  • Function execution (Operational)
View →
ResendMar 26, 2026 at 18:44 UTC

Email events delayed

Status: Resolved

Our systems have caught up and email events are processing normally. No emails were lost or delayed during this time.

Affected components
  • Email Events (Operational)
View →
InngestMar 26, 2026 at 14:55 UTC

Degraded function execution performance

Status: Resolved

After an extended monitoring period, we are resolving this incident. The system is full operational.

Affected components
  • Function execution (Operational)
View →
VercelresolvedMar 25, 2026 at 17:31 UTC

Elevated Dashboard Errors

Mar 25, 17:31 UTC
Resolved - This incident has been resolved.

Mar 25, 16:41 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Mar 25, 16:32 UTC
Investigating - We are currently investigating reports of elevated error rates on the Vercel Dashboard. Existing deployments and live traffic are not affected by this issue. We will share updates as they become available.

View →
ResendMar 24, 2026 at 21:23 UTC

Contacts are encountering errors when unsubscribing from the audience

Status: Resolved

We have resolved the underlying issue and service has been resumed.
View →
GitHubMar 24, 2026 at 20:56 UTC

Disruption with some GitHub services

Mar 24, 20:56 UTC
Resolved - This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

Mar 24, 20:38 UTC
Update - We are investigating elevated error rates affecting multiple GitHub services including Actions, Issues, Pull Requests, Webhooks, Codespaces, and login functionality. Some users may have experienced errors when accessing these features. Most services are now showing signs of recovery. We'll post another update by 21:00 UTC.

Mar 24, 20:23 UTC
Update - Issues is experiencing degraded performance. We are continuing to investigate.

Mar 24, 20:23 UTC
Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.

Mar 24, 20:20 UTC
Update - Webhooks is experiencing degraded performance. We are continuing to investigate.

Mar 24, 20:18 UTC
Investigating - We are investigating reports of degraded performance for Actions

View →
VercelresolvedMar 24, 2026 at 16:31 UTC

Elevated Errors Loading Deployments on Dashboard

Mar 24, 16:31 UTC
Resolved - This incident has been resolved.

Mar 24, 16:14 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Mar 24, 15:19 UTC
Investigating - We are currently investigating reports of elevated error rates loading the Deployment Overview on the Vercel Dashboard. Existing deployments and live traffic are not affected by this issue. We will share updates as they become available.

View →
VercelresolvedMar 24, 2026 at 05:01 UTC

Elevated errors creating Vercel Functions in sin1 (Singapore)

Mar 24, 05:01 UTC
Resolved - This incident has been resolved.

Mar 24, 04:40 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Mar 24, 03:43 UTC
Identified - We are observing elevated errors creating Vercel Functions in the `sin1` region. To mitigate errors creating deployments, we have temporarily disabled provisioning new Vercel Functions in this region. Existing deployments and live traffic are not affected by this issue. We will share updates as they become available.

View →
ResendMar 20, 2026 at 05:20 UTC

Delayed Email Events

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Email Events (Operational)
View →
ResendMar 20, 2026 at 04:02 UTC

Maintenance: Scheduled Maintenance on General API and Dashboard

Status: Complete

The maintenance was completed. We're working through some backlogged Email Events.

Affected components
  • Dashboard (Operational)
  • Email Events (Operational)
  • General API (Operational)
  • Webhooks (Operational)
View →
ResendMar 18, 2026 at 14:27 UTC

Deliverability Issues to Outlook Emails

Status: Resolved

The issue has been resolved, and sending to Microsoft-owned domains has been restored.
View →
VercelresolvedMar 16, 2026 at 18:44 UTC

Dubai region (dxb1) is unavailable and traffic is being re-routed

Mar 16, 18:44 UTC
Resolved - Due to ongoing issues in the dxb1 region, traffic and regional Vercel services are currently re-routed to bom1. We will share more details on recovery when they become available.

Mar 2, 18:23 UTC
Monitoring - We are monitoring the situation and continue to work toward restoring capacity in the dxb1 region. We will send further updates when new information is available.

Mar 2, 15:29 UTC
Identified - Due to operational issues in the dxb1 region, traffic is currently re-routed to bom1. Additionally, dxb1 is currently unavailable as a Function Region for new deployments.

If your existing deployments that use the dxb1 region are experiencing elevated function invocation errors, we strongly recommend switching to the nearest region (such as bom1) and redeploy until capacity is restored in dxb1. Deployments using multiple regions or failover regions are not affected since traffic is automatically routed to the nearest region based on the configured settings.

View →
VercelresolvedMar 13, 2026 at 23:15 UTC

Elevated Build Errors

Mar 13, 23:15 UTC
Resolved - This incident has been resolved.

Mar 13, 22:53 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Mar 13, 22:32 UTC
Investigating - We are currently investigating elevated error reports of deployment creation during the build phase. Existing deployments and live traffic are not affected by this issue. We will share updates as they become available.

View →
VercelresolvedMar 12, 2026 at 01:15 UTC

Elevated GitHub Deployment Failures

Mar 12, 01:15 UTC
Resolved - Between 01:15 and 05:45 UTC, customers might have experienced elevated GitHub deployment failures, including delays in triggering new deployments for commits and redeploy errors displaying the message "There is no GitHub account connected to this Vercel account". This was due to GitHub API degradations, which they have since resolved.

https://www.githubstatus.com/incidents/lw6j95nyw3py
https://www.githubstatus.com/incidents/02z04m335tvv

View →
InngestMar 10, 2026 at 02:31 UTC

Reduced throughput on function execution

Status: Resolved

The networking fix has been applied and all systems are operational. We have identified the root cause. Function execution has returned to normal. Any backlogs incurred during the incident will be executed.

Affected components
  • Function execution (Operational)
View →
VercelresolvedMar 6, 2026 at 21:38 UTC

Elevated Errors on Middleware Invocations

Mar 6, 21:38 UTC
Resolved - This incident has been resolved.

Mar 6, 21:20 UTC
Monitoring - A fix has been applied and we are seeing recovery for affected deployments. We are continuing to monitor.

Mar 6, 19:18 UTC
Update - We are applying a fix for the deployments experiencing elevated errors.

We continue to recommend redeploying if you are seeing errors on deployments created between 11:20 UTC and 15:14 UTC. Deployments created outside of this window are unaffected and no action is required. Additionally, deployments with middleware on the Node runtime are unaffected.

Mar 6, 15:25 UTC
Identified - Some deployments created between 11:20 UTC and 15:14 UTC with Edge Middleware may be seeing elevated errors. Deployments created outside of this time window are unaffected. If you are experiencing issues, we recommend redeploying.

View →
VercelresolvedMar 6, 2026 at 21:22 UTC

Elevated Latency on Queue Messages, Increased Message Retry Latency in iad1

Mar 6, 21:22 UTC
Resolved - This incident has been resolved.

Mar 6, 21:10 UTC
Monitoring - A fix has been implemented and we are monitoring the results.

Mar 6, 21:00 UTC
Identified - The issue has been identified and a fix is being implemented.

Mar 6, 20:53 UTC
Investigating - We are investigating reports of elevated queue message latency and message retry latency in iad1, as well as elevated sleep times and increased step retries in Workflow. We will provide additional updates as they become available.

View →
ResendMar 5, 2026 at 23:55 UTC

Email events delayed

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Email Events (Operational)
View →
InngestMar 3, 2026 at 23:44 UTC

Function execution delayed

Status: Resolved

The incident is now resolved and the system is full operational after monitoring.

Affected components
  • Function execution (Operational)
View →
VercelresolvedMar 3, 2026 at 21:10 UTC

Delays Loading Observability, Usage, Analytics, and Speed Insights Data

Mar 3, 21:10 UTC
Resolved - This incident has been resolved.

Mar 3, 19:14 UTC
Monitoring - A fix has been implemented and we are seeing recovery for data loading and ingestion across services. We are continuing to monitor and will provide additional updates as they become available.

Mar 3, 17:10 UTC
Investigating - Dashboard pages that use Observability data, including Observability, Speed Insights, Web Analytics, Usage, Firewall, and Activity, are experiencing delays while loading data. These pages are also experiencing delays ingesting new data. We are investigating this issue and will provide additional updates as they become available.

View →
ResendMar 3, 2026 at 07:47 UTC

Email sending delay

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
View →
VercelresolvedMar 2, 2026 at 15:20 UTC

Elevated deployment and function invocations failures

Mar 2, 15:20 UTC
Resolved - This incident has been resolved.

Mar 2, 14:59 UTC
Monitoring - We have rolled out a second mitigation for elevated Build errors and are seeing recovery. All builds are now excluding the Dubai region (dxb1) from their deployment targets as a temporary measure. We will provide additional updates as they become available.

Mar 2, 13:00 UTC
Update - We have rolled out a first mitigation for elevated Build errors. Builds that use Middleware are now excluding the Dubai region (dxb1) from their deployment targets as a temporary measure, and should complete successfully again. We are now working on a mitigation for Builds that are using Edge Functions.

Mar 2, 11:59 UTC
Update - We are currently deploying a mitigation for elevated Build errors. Builds that use Middleware or Edge Functions will exclude the Dubai region (dxb1) from their deployment targets as a temporary measure. We will provide additional updates as they become available.

Mar 2, 10:50 UTC
Update - We are still seeing elevated errors in Builds in all regions, because Middleware and Edge Functions may be deployed globally. Builds that don't use Middleware and Edge Functions are not impacted. We are continuing to work on a fix for this issue.

Mar 2, 08:43 UTC
Update - The dxb1 Edge traffic is currently being rerouted to the nearest Edge region (bom1) to mitigate the impact. We will provide additional updates as they become available.

Mar 2, 06:24 UTC
Update - We have rolled out mitigations and are seeing recovery. If you are still seeing build failures and are using Dubai (dxb1) as your primary Vercel Functions region, you can switch to another region as a workaround.

Mar 2, 06:06 UTC
Identified - Starting from 5:00 am UTC, we have started seeing failures to deploy and invoke functions in Dubai region (dxb1). Deployments with Middleware Functions are also impacted in all regions, because Middleware Functions are deployed globally for production deployments. Our team is actively investigating the issue.

View →
VercelMar 2, 2026 at 12:08 UTC

Degraded Logs and Traces in Dubai region (dxb1)

Mar 2, 12:08 UTC
Resolved - This incident was resolved.

Mar 2, 08:49 UTC
Monitoring - The dxb1 Edge traffic is currently being rerouted to the nearest Edge region (bom1) to mitigate the impact. We will provide additional updates as they become available.

Mar 2, 08:01 UTC
Identified - We are continuing to work on a fix for this issue.

Mar 2, 06:54 UTC
Investigating - We are currently investigating issues collecting Logs and Traces in Dubai (dxb1). We will share updates as they become available.

View →
ResendMar 2, 2026 at 09:15 UTC

Increased latency in email sending

Status: Resolved

Performance is back to normal.

Affected components
  • Email Events (Operational)
  • Webhooks (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Website (Operational)
  • Dashboard (Operational)
View →
VercelMar 2, 2026 at 00:22 UTC

Deployments are failing for customers using Static IPs in Dubai region (dxb1)

Mar 2, 00:22 UTC
Resolved - A fix has been implemented and new deployments are no longer failing.

Mar 1, 21:17 UTC
Identified - Deployments for customers that use Static IPs are failing when deploying functions to the Dubai region. Customers can remove the Dubai region from the Static IPs configuration to avoid deployment failures. Our team is actively investigating the issue. Deployments to other regions are not affected.

View →
ResendFeb 27, 2026 at 18:30 UTC

Increase latency in email sending

Status: Resolved

Performance is back to normal.

Affected components
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • Single Email (Operational)
View →
InngestFeb 26, 2026 at 02:07 UTC

Elevated latency for function execution

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
ResendFeb 25, 2026 at 14:16 UTC

Delay in email sending

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • Dashboard (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
View →
ResendFeb 24, 2026 at 20:59 UTC

Delay in email sending

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Batch Emails (Operational)
  • Single Email (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
View →
InngestFeb 19, 2026 at 18:47 UTC

Degraded dashboard availability

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Inngest Dashboard (Operational)
View →
ResendFeb 17, 2026 at 00:33 UTC

Errors with Resend Dashboard and Email Latency

Status: Resolved

This incident has been resolved. Email sending and dashboard performance are both operating normally. Thank you for your patience while we worked through this.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
View →
ResendFeb 16, 2026 at 23:32 UTC

Errors Loading Resend and Email Sending Delays

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
View →
ResendFeb 13, 2026 at 02:35 UTC

Intermittent Issues with SMTP

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • SMTP (Operational)
View →
ResendFeb 12, 2026 at 14:36 UTC

Increased latency in Broadast sending

Status: Resolved

Service has been resumed.

Affected components
  • Broadcast Emails (Operational)
View →
InngestFeb 10, 2026 at 15:57 UTC

Delays on some function execution

Status: Resolved

The incident is now resolved and the system is full operational.

Affected components
  • Function execution (Operational)
View →
InngestFeb 9, 2026 at 22:50 UTC

Delayed function execution

Status: Resolved

After an extended monitoring period, function execution has returned to normal rates across all queue shards. During this issue, only a subset of users were affected on part of our infrastructure. Our infrastructure team is in the midst of rolling out additional system capacity going forward.

Affected components
  • Function execution (Operational)
View →
InngestFeb 9, 2026 at 07:51 UTC

Delayed function run status

Status: Resolved

Runs, traces and events data are all caught up from their temporary backlog. The dashboard metrics are all being processed with no backlog.

Affected components
  • Observability (Operational)
View →
ResendFeb 6, 2026 at 05:52 UTC

Slow Broadcast Queue Processing

Status: Resolved

Broadcast sending has returned to normal speed. Queue times are back to expected levels. Thank you for your patience.

Affected components
  • Broadcast Emails (Operational)
View →
ResendFeb 5, 2026 at 17:10 UTC

Delay in email sending

Status: Resolved

It was a false alert. It was caused by a combination of an ongoing incident with the third‑party observability service we use and some internal, isolated infrastructure experiments. Emails have been sent smoothly and there have been no delays.

Affected components
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Batch Emails (Operational)
  • Webhooks (Operational)
View →
ResendFeb 4, 2026 at 17:48 UTC

Errors on API

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • Webhooks (Operational)
  • Website (Operational)
  • General API (Operational)
View →
InngestFeb 3, 2026 at 23:36 UTC

Subset of customers experiencing function execution delays

Status: Resolved

System latency for function execution has returned to normal levels for the affected users. The incident has been resolved. The cause of the incident was due to increased load causing congestion. We applied changes to the system to reduce congestion, resulting in increasing throughput. We also re-distributed some affected users in an effort to mitigate impact. Our team's planned to roll out new infrastructure in the coming weeks and is accelerating that plan to aim to roll it out later this week to increase overall capacity.

Affected components
  • Function execution (Operational)
View →
ResendFeb 3, 2026 at 23:21 UTC

Latency increased for all background actions

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • General API (Operational)
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
View →
ResendFeb 3, 2026 at 13:48 UTC

Increased latency

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Batch Emails (Operational)
  • General API (Operational)
  • Webhooks (Operational)
  • Single Email (Operational)
  • Dashboard (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
View →
ResendFeb 3, 2026 at 08:31 UTC

Increased latency

Status: Resolved

We have resolved the underlying issue and service has been resumed. All emails that have queued have been sent.

Affected components
  • General API (Operational)
  • Webhooks (Operational)
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Batch Emails (Operational)
  • Website (Operational)
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
View →
ResendFeb 2, 2026 at 17:57 UTC

Email sending delayed

Status: Resolved

We have resolved the underlying issue and service has been resumed.

Affected components
  • Dashboard (Operational)
  • Broadcast Emails (Operational)
  • Batch Emails (Operational)
  • General API (Operational)
  • Website (Operational)
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Webhooks (Operational)
View →
ResendJan 30, 2026 at 11:58 UTC

Resend team creation error

Status: Resolved

The underlying issue has been identified and resolved, and services have been fully restored. Thanks for your patience while we worked through this.
View →
InngestJan 28, 2026 at 01:49 UTC

Error publishing events with metadata to Event API

Status: Resolved

We shipped a fix earlier today at 20:40 UTC (Jan 27) that has resolved the issue after an extended monitoring window to confirm the issue did not return. The bug itself affecting some user requests was an issue due to a large "baggage" header which was beyond the limit of what the Event API's pubsub event stream can handle. Baggage headers are used for sending extra context within a request like open telemetry or APM tracing data. Some requests contained more than 1024 bytes which caused this issue. The fix applied earlier today gracefully now gracefully handles the situation where there are large baggage headers. This issue will not surface again.

Affected components
  • Event API (Operational)
View →
ResendJan 27, 2026 at 17:50 UTC

Delays in email sending

Status: Resolved

Our systems have caught up with the delayed emails and are healthy again. No emails were lost during the period of degraded performance.

Affected components
  • Batch Emails (Operational)
  • Single Email (Operational)
  • Email Events (Operational)
  • SMTP (Operational)
  • Broadcast Emails (Operational)
View →
ResendJan 27, 2026 at 12:45 UTC

Internal server error returned from the API

Status: Resolved

We have resolved the underlying issue and confirmed the service has been resumed.

Affected components
  • General API (Operational)
View →
ResendJan 23, 2026 at 05:48 UTC

Microsoft 365 Incident

Status: Resolved

Microsoft is reporting that email is returning to normal sending, with its most recent update stating, "We have a high level of confidence that the incident is largely resolved." From our own metrics, we're also seeing "Delivery Delayed" events reduce, with deliveries now happening upon retry.

Affected components
  • Email Events (Operational)
View →