Back to overview
Resolved

Partial Service Disruption

Jul 30, 2026 at 12:03am UTC
Affected services
Global Network

Resolved
Jul 30, 2026 at 12:43am UTC

Resolved

Recovery has been completed, and the incident has been fully resolved.
Service was progressively restored through controlled recovery procedures following successful infrastructure isolation, validation, and capacity rebalancing. Extended health verification has confirmed stable AI model serving operations across the affected region.

We appreciate your patience while our Engineering, Infrastructure, and Site Reliability teams executed the recovery. A comprehensive post-incident review is now underway to identify further opportunities to strengthen operational resilience.

Updated
Jul 30, 2026 at 12:34am UTC

Monitoring

Mitigation has been successfully completed, and service has been restored across the affected AI Model Serving infrastructure.

We are currently performing controlled validation and continuously monitoring network health to verify service stability, workload distribution, and operational integrity before declaring the incident fully resolved.

Updated
Jul 30, 2026 at 12:29am UTC

The affected GPU serving infrastructure has been successfully isolated to contain the impact and prevent further service degradation.

Traffic has been redirected to healthy serving capacity, and our engineering teams continue executing recovery procedures while validating AI model serving operations.

Updated
Jul 30, 2026 at 12:22am UTC

As a precaution, we have activated additional traffic safeguarding and infrastructure protection measures across our global AI serving platform while recovery efforts continue.

Updated
Jul 30, 2026 at 12:19am UTC

We have initiated mitigation and recovery procedures as our on-call engineers continue working to restore affected services. Recovery efforts are progressing, and service availability is gradually improving across impacted infrastructure.

Updated
Jul 30, 2026 at 12:11am UTC

Identified
We have identified the issue affecting our AI Model Serving service in the Bahrain region. The disruption is due to a regional infrastructure impact associated with an ongoing political event, affecting a portion of our distributed GPU capacity in the region.

Some AI models may experience increased latency, intermittent errors, or temporary unavailability while recovery efforts are underway. Our engineering team is actively mitigating the impact and working to restore full service as quickly as possible.

We will continue to provide updates as more information becomes available.

Created
Jul 30, 2026 at 12:03am UTC

We are aware of an ongoing issue impacting our distributed GPU infrastructure in the Bahrain region. As a result, some AI models may experience service disruptions, increased latency, or temporary unavailability.