Blog · 5 min read ·
ShareAzure gateway outage: your VPN backup shared the failure
Between 20:30 UTC on 30 September 2026 and 02:15 UTC on 01 October 2026, Microsoft says, "a subset of customers using gateway services in multiple regions experienced degraded or interrupted network connectivity." The incident is Tracking ID 7Q30-010 in Microsoft's Azure status history, titled "Mitigated- Multiple services experiencing connectivity issues in multiple regions." Two of the services on its impacted list are the ones many mid-market hybrid designs treat as primary and backup: Azure ExpressRoute Gateway and Azure VPN Gateway.
What Microsoft's entry says
Microsoft lists five impacted services: Azure ExpressRoute Gateway, Azure Firewall, Azure Application Gateway and Web Application Firewall, Azure VPN Gateway, and Azure VMware Solution. It says impacted customers "may have observed gateways failing to load in the Azure Portal, along with failures or delays in network management operations."
On cause, Microsoft says "a recent change to a regional gateway management service triggered a higher-than-expected load when an unrelated operating system servicing maintenance proceeded gradually through multiple regions. Normally, this gateway management service would auto scale as needed." The fix was to revert: "We reverted the contributing regional gateway manager change, which reduced gateway manager load and allowed affected services to recover."
The timeline, all UTC:
- 20:30, September 30: customer impact began.
- 21:29: Microsoft began investigating connectivity issues affecting ExpressRoute Gateway in UK South.
- 22:27: Microsoft identified that multiple regions were affected.
- 23:05: Microsoft correlated the problem with operating system servicing activity and paused further servicing.
- 01:36, October 1: recovery progressed across most affected regions, while configuration changes were applied to remaining regions, "including France Central, North Europe, Southeast Asia, UK South, and UK West."
- 02:15: Microsoft confirmed the issue was mitigated.
Two cautions on scope. First, Microsoft's entry carries a note that "Impacted region(s) within a previous communication has been corrected," so confirm the region list against the live entry rather than a screenshot or a summary, ours included. Second, Microsoft says the times above "represent the full incident duration, so are not specific to any individual customer." Only your own telemetry can say whether your gateways were touched.
Redundant path, shared control plane
A common mid-market hybrid design: ExpressRoute carries production traffic to Azure, and a Site-to-Site VPN through an Azure VPN Gateway sits behind it as the failover. It is sensible against the failures it was built for: a carrier fault, a cut fiber, a misbehaving circuit. A second path through a different medium handles those.
This incident was different in kind. Microsoft's own account ties the impact to a regional gateway management service, and both gateway types are on its list. From the entry, our reading is that the two shared a dependency on Azure's management layer for gateways, even though they are different products with different transport. We can't see inside Microsoft's architecture, and the PIR may refine this. The pattern is the point: a backup on the same management service as the primary is redundant in transport and shared in control.
The symptom Microsoft describes is also worth reading carefully. Gateways failing to load in the portal, and "failures or delays in network management operations," means that a failover plan whose step one is "log in and change a route or a gateway setting" may stall at step one. You can't count on reconfiguring Azure networking mid-incident, because the reconfiguration path may be part of what is degraded.
We are not claiming that any customer's ExpressRoute and VPN failed together. Microsoft doesn't say which customers were affected. The claim is narrower: on Microsoft's own list, primary and backup were both named in one incident.
A checklist for this week
- Draw the real dependency map. For each Azure region you use, list every path in and out: ExpressRoute, VPN Gateway, Azure Firewall, Application Gateway or WAF, and Azure VMware Solution if you run it. Mark which ones Microsoft listed. Anything listed that your failover relies on is a shared dependency.
- Find the failover that needs no Azure change. The best backup path comes up without you touching Azure during the incident: pre-built, pre-tested, failing over on its own. If yours needs a portal click, treat that as a gap.
- Price a path that does not use an Azure gateway. Whether it is worth the cost depends on what the link carries and what an hour of it being down costs you.
- Turn on Azure Service Health alerts and route them somewhere a person will see them at 3 a.m. Microsoft's entry itself tells customers to "configure and maintain Azure Service Health alerts," and says that is also how you get notified when the PIR is published.
- Check your own telemetry for the window. Pull gateway and tunnel logs for 20:30 UTC September 30 through 02:15 UTC October 1. Microsoft's times are the full incident duration; yours may be shorter, or zero.
- Record the incident ID and your times if you will claim a credit. Write down 7Q30-010, the UTC window, the resources involved, and your own evidence of impact while the logs still exist. We did not retrieve Microsoft's SLA page for this post, so we quote no credit amounts. Read your own agreement for what counts as downtime, what is eligible, and the claim window.
- Test the failover on purpose. In a maintenance window, pull the primary and watch what actually happens, including the steps a person has to do.
For the adjacent lesson about redundancy and recovery, see our post on why redundancy is not a backup, and for the contractual side of provider outages, what Microsoft's service-credit policy actually covers.
Where we fit
We don't resell hardware or software. When a project needs procurement, we recommend the right products for your environment, work with the vendor of your choice on quoting, and handle install, configuration, and integration. No margin in either direction. For this problem, that means mapping your real dependencies, testing the failover, and writing the runbook.
Sources
The 30-second version
Microsoft says a subset of customers using gateway services in multiple regions had degraded or interrupted connectivity from 20:30 UTC September 30 to 02:15 UTC October 1, 2026, with ExpressRoute Gateway and VPN Gateway both on its impacted list. It fixed the problem by reverting a change to a regional gateway management service. Its Post Incident Review is pending, and we will update this post. If your VPN is the backup for your ExpressRoute, check whether the two share a dependency you haven't named, and whether your failover needs an Azure change to work. If you want a senior engineer to map that for you, the project intake form takes about three minutes.
Pro IT NW helps regulated mid-market organizations design, implement, and test backup and disaster recovery for Microsoft 365, Azure, AWS, on-premises, and hybrid environments. Vendor-neutral, labor-only. This post reflects Microsoft's Azure status history entry as of October 5, 2026; the Post Incident Review was not yet published.
Questions we get asked
- What did Microsoft say happened in the September 30, 2026 Azure gateway incident?
- Microsoft's Azure status history entry (Tracking ID 7Q30-010) says that between 20:30 UTC on 30 September 2026 and 02:15 UTC on 01 October 2026, a subset of customers using gateway services in multiple regions experienced degraded or interrupted network connectivity. It attributes the incident to a recent change to a regional gateway management service that triggered higher-than-expected load while unrelated operating system servicing maintenance proceeded through multiple regions. Microsoft reverted the change and the affected services recovered.
- Did the incident affect both ExpressRoute and VPN Gateway?
- Both appear on Microsoft's list of impacted services: Azure ExpressRoute Gateway, Azure Firewall, Azure Application Gateway and Web Application Firewall, Azure VPN Gateway, and Azure VMware Solution. Microsoft describes the impact as affecting a subset of customers, so appearing on the list does not mean any particular customer's gateways failed.
- Has Microsoft published a Post Incident Review?
- Not as of this writing. Microsoft's entry says that once its internal retrospective is complete, generally within 14 days, it will publish a Post Incident Review (PIR) to all impacted customers. We will update this post when the PIR is available.
Related service
Backup & disaster recovery implementationWritten by the team at Pro IT NW · Senior-led Microsoft project consultancy · Seattle and the Pacific Northwest, delivered USA-wide.