The Time it Wasn't DNS: Modern Incident Analysis and the Azure Global WAN Outage

2026-07-2825 min read

Every organization has a favourite outage explanation. Sometimes it is DNS. Sometimes it is "the engineer didn't follow the runbook". Sean Klein's argument is that both are the same failure of analysis: a complex system produced a complex failure, and the organization compressed it into a story simple enough to repeat to an executive. The compression feels like understanding, but it produces repair items that fix problems which never existed while leaving the real systemic weaknesses in place.

Klein is a Principal Technical Program Manager at Microsoft, where he leads the Production Livesite Review program for Azure. He describes his job, only half jokingly, as turning outages into twenty-page Word documents and delivering them to senior leadership. What distinguishes his practice — which he calls modern incident analysis — from ITIL problem management and its derivatives is not the document but the method: for a certain class of incidents, Azure goes much deeper than the Five Whys. He has worked on post-incident activity for roughly two decades, previously at Salesforce and in private consulting, and is a member of the Resilience in Software Foundation. This 43-minute, 59-second talk was recorded at QCon San Francisco 2025 and published by InfoQ on June 23, 2026.

These notes report what Klein presented and clearly mark the small amount of background context added for readers unfamiliar with routing protocols.

What You Will Learn

  • Why the "simple story" of operator error is attractive, and what it costs an organization when it becomes the accepted narrative.
  • The full contributing-factor chain behind Azure's January 25, 2023 global WAN outage, spanning two months of preceding decisions.
  • How a contributing-factor diagram is used as an interviewing and alignment technique rather than as an incident specification.
  • How to read such a diagram for tactical repairs versus systemic repairs.
  • Why safety science treats Standard Operating Procedures as artifacts that humans interpret rather than instructions humans execute.
  • How to introduce blameless analysis inside an organization whose culture is not yet blameless, and where to find the source material.

Azure's Incident Vocabulary

Klein opens with terminology, because the words carry organizational meaning that outsiders routinely misread.

Term Meaning inside Azure
DRI / OCE Directly Responsible Individual, or on-call engineer — the person the pager wakes up. Klein uses them interchangeably.
MOP / SOP Method of Procedure and Standard Operating Procedure. Technically a MOP is an instance or child of an SOP; for this talk they are interchangeable, and Klein notes his own leadership used them interchangeably in the incident artifacts.
Incident What many organizations would call an alert. A monitoring threshold is breached, an incident fires and lands in a queue.
Outage A declared state, raised when customer impact is confirmed or imminent. Declaring an outage can invoke the centralized incident management function.

The distinction matters practically. An Azure engineer saying they have 500 incidents in their queue is describing something bad, but not 500 outages.

Outages carry a severity scale. Severity 1 is the largest impact tier, reserved for outages breaching a promised containment zone: multi-service, a full availability zone or several, multi-region, or global. Above that sits a Sev 0, which Klein describes as a Sev 1 plus a great deal of additional machinery. It does not change the incident response — nobody responds harder to a Sev 0 — but it triggers extra workstreams for internal and external communications and commitments. Customers cannot see this scale directly. Klein's practical indicator is that if the outage is on Downdetector, CNN, Reuters, or ZDNET, or if Mary Jo Foley is tweeting about it, it is probably a Sev 0. These are the named outages, the ones still discussed three years later. Say "the global WAN outage" inside Microsoft and everyone knows which one you mean.

That outage occurred on January 25, 2023. Microsoft was effectively hard down globally for about an hour and forty minutes. The WAN is how customers reach Azure, and it was unavailable. It is also used by other Microsoft clouds — Klein names Office/M365 and Xbox — so any Microsoft customer at that time was affected.

Why the Simple Story Forms

Klein makes a point that the narrative starts forming during the outage, before mitigation. In this case it was "a change gone wrong", which, as he concedes, is not incorrect. The problem is what happens after the dust settles.

Giant outages pull in people who do not normally handle outages. Klein's team lives in outages daily; most of the company does not. Senior leaders, customer- facing staff, and communications people suddenly have to explain the event to furious customers, to a board of directors, or to external media. They are doing exactly what an on-call engineer does when joining a bridge: synthesizing partial data, reading logs and alert patterns, forming theories, and simplifying until the situation is tractable. The difference is that they are doing it hours later, from chat logs and second-hand summaries, and under pressure to produce something communicable.

Klein showed anonymized screenshots of real internal comments, Word document annotations, and Teams messages. He anonymized them to keep the analysis blameless and explicitly not to shame anyone, because this happens in every organization. His diagnosis is that it stems from a universal human need to make a simple story out of a complex problem — part of the ordinary sense-making process that engineers and PMs both rely on. Left unchecked, that simplification becomes the official story.

His thesis in one line: it is never a simple story. It is never DNS; even when it is DNS, it is never just DNS. Hence the title — this time it was BGP, and it is also never just BGP. It is never just operator error, and never just an engineer who failed to follow the SOP.

The One Why

To make the point concrete, Klein diagrams the simple story: a single node, one why instead of five, which he dryly calls an 80% efficiency increase. The story is attractive because it is easy to tell, easy to understand, and it is a canonical outage type. It even generates its repair items automatically, without needing a dropdown: punish the engineer who made the change — name them, or in an extreme case fire them — and for everyone still employed, increase the mandatoriness or the frequency of the mandatory training.

The story arrives tied with a bow, and it is immediately accepted. That is the danger. If the organization acts on it, it fixes problems that did not exist while causing real cultural harm — punishing someone for something they did not actually do, or without understanding the factors that shaped their decision, and then telling everyone else they need more training. None of the underlying issues are touched.

The Contributing-Factor Diagram

Klein's alternative is a large diagram of contributing factors, deliberately shown at a size the audience cannot read, because the illegibility is the point. He then zooms in progressively.

Two caveats govern how the artifact should be understood, and Klein repeats them several times. First, the diagram is not a specification of the incident and is not scientific truth; it is whiteboarding and brainstorming used while developing the narrative. Second, its primary purpose is as an interviewing and alignment technique. Klein builds it collaboratively with incident stakeholders, service owners, engineering teams, and the on-call engineer who executed the change. He works with one group to produce part of the diagram, then takes it to another group who tell him a portion is wrong. That friction is the method working. The goal is to break participants out of the root-cause construct — the belief that asking why five times magically yields the cause, and that the cause is the thing to fix.

Each node is a contributing factor: an event or condition that had to be true, or had to not be true, for the customer to experience the impact in the way they did. Stepping into hindsight and counterfactuals, the corollary is that removing any single node would have prevented the impact. Klein finds that a far more attractive story than firing an engineer and mandating more training, precisely because it offers many places to intervene.

What Actually Happened on January 25, 2023

The impact came in three distinct periods: the roughly hour-and-forty-minute global hard-down, and then two long-tail recoveries in different places — West India, and a North American point of presence around Chicago — each for different reasons.

At the centre of the diagram sits the simple story, because it did happen: an engineer ran a command on a router and the world went dark for an hour and forty minutes. Understanding it requires going back two months.

The Build-Out and the Re-IP

The target router was a new role on the WAN, part of a network expansion increasing both the size and the speed of the global WAN. As a WAN role, it was provisioned with an external IP address, which is how wide area networks peer to the internet. The cohort was not a single device; it was twelve router pairs.

During the build-out, the architecture changed. These routers needed to become part of Microsoft's Software-Defined Wide Area Network (SWAN); Klein points readers to Microsoft's published SWAN material rather than covering it in the talk. A nuance of that design is that the routers required an internal IP, so already-built routers had to be re-IP'd. Klein notes this is a reliable way to identify the network engineers in an audience, because they start squirming — re-IP is not a typically low-risk operation.

The SOP That Skipped Governance

Because the role was new, no SOP existed for the re-IP, so one had to be developed. Azure network changes carry a high degree of governance, and that extends to SOP governance: creating or modifying an SOP normally passes through an extensive change review process. One of its five steps is emulation. The entire Microsoft WAN is virtualized inside Azure VMs, so a change can be applied virtually and observed before it touches production. That is one step of five required simply to update an SOP that will later drive a production change.

Critically, these routers were never considered production at any point during the incident. They were not serving customer traffic and were not communicating with other routers. They were, however, connected to the IGP backplane — Klein flags this explicitly as foreshadowing.

Because the work was classified as non-production, it fell into an area where practices around SOP change processes were inconsistent. The SOP was created and later modified outside the governance process: no peer review, no CAB review, no testing in the emulated environment.

The Command Gets Added

The SOP was used successfully on a couple of earlier routers in the cohort. On the most recent execution before the outage, engineers hit a problem: after the re-IP, the link state packets were stale. Resolving that normally requires deleting the database inside the router holding the routing configurations. Because the role was new, the team asked the router manufacturer for guidance, and the vendor supplied a command. That command was added to the SOP — again, outside the governance process one would expect.

(Background for readers new to routing: a link-state protocol has each router flood descriptions of its own links to its neighbours, and every router builds a map of the network from those advertisements. "Stale link state packets" means that map has diverged from reality, which is why clearing the local database is a plausible remedy.)

Then the calendar intervened. The holiday period arrived, with targeted change freezes protecting retail customers and much of the staff and customer base away. Large change work paused for a few months.

Execution Day

On January 25, it was time to run the operation on the next router. The engineer followed the SOP: it was attached to the change ticket, they opened it, read it, and executed — unaware that the SOP had been modified since they last saw it.

Klein uses this to make a safety-science argument, referencing a point made in another talk at the conference about how runbooks should be treated. Tech is famous for believing an SOP should be infallible: you run it, and deviating from it is a problem. Safety science and resilience engineering take the opposite view. Humans interpret SOPs and make assumptions and decisions constantly, and no SOP can ever be perfect. Klein says explicitly that Azure wants humans interpreting the SOP — wants them to know when what they are doing is high risk, to apply scrutiny, and to escalate or question a command they think might be unsafe.

Which raises the obvious question: why did this engineer not question a command that had been added out of band? Two elements of their mental model explain it, and both are properties of the system, not of the person.

The multi-vendor fleet. The Microsoft WAN uses three different router manufacturers across its roles, which means three different operating systems and multiple versions of each are represented in the network. Microsoft does this for several reasons, one being supply chain de-risking. The engineer did see the command and was familiar with it — on two of the three operating systems in the fleet, that command is locally scoped and safe to run. On the third, it affects adjacency. When run there, its scope was effectively global.

The expectation of a guardrail. The engineer also expected that an unsafe command would be blocked at the AAA (authentication, authorization, and accounting) layer of the network. It should have been. When a new role is onboarded, Azure performs a large audit identifying every command capable of more-than-local impact on a device and blocks it. In this case the audit was in progress but had not yet reached the full command audit — purely an order-of-operations timing problem in the onboarding sequence.

So the engineer ran the command confidently. So confidently that 33 minutes later they ran it on the other router of the pair.

On a production change there would have been a listening period — implement, then spend an enforced hour or so watching key health indicators before proceeding. That step was not in this SOP, because this was not classified as a production SOP change, and from the engineer's perspective they were not making a production change. They had no visibility into the impact of the first command. The network was in fact almost healed from the first execution when the second one landed.

The Cascade

The combination produced two cascading events. Klein gives a deliberately simplified breakdown rather than a deep dive into routing theory: routers use IGP (Interior Gateway Protocol) to manage internal communications, and BGP (Border Gateway Protocol) to manage peering with external internet peers. The command resets the IGP routing table on the local device — except here, on the entire network.

The whole Microsoft WAN began recomputing its internal connectivity. That ran for about 30 minutes, healed slightly, and then the second execution triggered it again. That second round is what pushed BGP into recomputation. Given the sheer size of the Azure WAN, roughly 15 million routes had to be recomputed, taking about an hour and forty-five minutes to complete.

The two long tails came from separate causes. During the IGP reconvergence, three devices actually failed; when the WAN came back up, those three were degraded, due to defects in the devices themselves. Compounding this, during the outage Microsoft had paused its auto-detection, auto-rerouting, and auto-recovery system, because with everything down it was contributing to the problem rather than helping. Traffic in those regions was not restored until that system was re-enabled, the degraded routers detected, and then manually ejected.

Architecture And Data Flow

The following diagram summarizes the contributing-factor chain Klein walked through. It is a simplified rendering of his whiteboarding artifact, not a specification of the incident.

flowchart TD
    A["WAN expansion: new router role,
12 router pairs, external IPs"] --> B["Architecture change to SWAN
requires internal IPs"] B --> C["Re-IP SOP must be created
(new role, no existing SOP)"] C --> D["Non-production classification:
inconsistent SOP change practices"] D --> E["SOP created and modified outside governance:
no peer review, no CAB, no emulation test"] F["Earlier execution hits
stale link state packets"] --> G["Vendor supplies command
to clear routing database"] G --> E H["New role onboarding audit
had not reached command audit"] --> I["Unsafe command not blocked
at AAA layer"] J["Three router vendors in fleet;
command is locally scoped on two OSes"] --> K["Engineer's mental model:
command is safe and local"] E --> K I --> K K --> L["Jan 25: command executed
on router 1"] M["No listening period in SOP
(not a production change)"] --> N["Command executed on router 2
33 minutes later"] L --> N L --> O["IGP reconvergence
across entire WAN, ~30 min"] N --> P["BGP reconvergence:
~15 million routes, ~1h45m"] O --> P P --> Q["Global hard down
~1h40m"] O --> R["3 devices fail during
reconvergence (device defects)"] S["Auto-detect / reroute / recover
paused during outage"] --> T["Long-tail impact:
West India, Chicago PoP"] R --> T

The Root Cause, Stated Honestly

Klein returns to the one-why diagram and answers the question people keep asking him. Asked for the root cause, he gives it as a single sentence, roughly: an engineer executed what was understood to be a low-risk, locally scoped planned change to a non-production router pair, following an official Standard Operating Procedure that had previously been updated by senior engineers to incorporate new guidance from the router manufacturer addressing issues encountered during an earlier execution of the same operation, with a mental model informed by extensive experience with other router manufacturers and a belief that any potentially unsafe, globally scoped command would be blocked by the AAA system.

His follow-up questions are the actual argument: what repairs do you write for that? Who do you fire for that? What, in that sentence, is the root cause?

Reading the Diagram for Repairs

The diagram's practical value is that it gives you a topology to read.

Node degree signals systemic issues. Nodes with many arrows entering or leaving are good indicators of a systemic or key issue, or something otherwise worth deep interest.

Distance from impact separates tactical from thematic. Contributing factors closest to the impact tend to map to tactical repair items — the things most likely to prevent this exact outage from recurring. Azure did exactly this: the relationship where IGP recomputation forces a BGP recompute has various buttons and levers available to make BGP more robust and less sensitive to what Klein calls IGP shenanigans. They pulled those levers days later, hardening the whole network against that failure class.

Factors further from the impact are the systemic and thematic ones. SOP change governance is Klein's example: the useful question is how many outages that weakness had already caused or nearly caused. If your goal is to add resilient behaviours to systems and to the organization rather than to prevent one specific repeat, that is where to focus.

Mental models are a legitimate repair target. Klein criticized training earlier, but qualifies it: training is good when it is right-sized and meets engineers where they are. Rather than making existing training mandatory with twice-yearly attestation, Azure now studies this incident as part of onboarding training for the WAN team and the core networking team. Part of the intent is to instil respect for a complex system that fails in complex and often unpredictable ways, even while running an SOP.

Transparency and Public Sources

Klein emphasizes that almost nothing in the talk was new disclosure — it had already appeared in Microsoft's post-incident reports. For many outages Azure runs Azure Incident Retrospectives, live Q&A panels with engineering leaders. Anyone can register, customer or not, and submit questions to the panel during the live event. Public PIRs are published at azure.status.microsoft, with the recording of the live session at the top.

Trade-offs And Limitations

This method does not scale to every incident. Klein is explicit that this is not how every Azure incident is postmortemed. The organization still pushes metrics and dashboards to leadership as fast as possible, and will continue to. Deep analysis is reserved for a particular class of incidents.

It is slow, and should not be time-boxed. A typical analysis of his takes about a month. Klein argues it should not be time-boxed against someone's reporting commitments. His deep work melds into the fast-reporting track rather than replacing it: for a big incident he is typically involved in the first few days, helping get communications to the right people and ensuring what is said is truthful and sufficiently informative.

The diagram is fragile if mistaken for truth. Klein repeatedly disclaims it as brainstorming, not specification. Presenting a contributing-factor diagram as authoritative invites exactly the false confidence the method exists to dissolve — and the counterfactual framing ("remove any node and there is no impact") is an analytical device, not a claim about how the system would actually have behaved.

Culture change cannot be a prerequisite. Asked how to shift the narrative for the general public, given how aviation has largely accepted that you do not blame the pilot, Klein's answer is that if the first step of implementing your program is changing your organization's culture, the program will fail. You have to work within the culture you have and change it gradually from the grassroots — sometimes keeping the work under the radar. He recommends borrowing examples from aviation, healthcare, and other safety-critical fields: when a plane crashes, nobody's goal is getting an MTTR metric to the president as fast as possible; everyone waits a year for the NTSB report because the objective is to learn.

The simple narrative forms whether or not anyone drives it. Asked whether he had to actively facilitate against the simplistic story internally, Klein said yes. He compares the narrative to a chord in music — it strikes as familiar even if you have never heard it before, so "human error, was he fired?" assembles itself without anyone deliberately constructing it. Countering it requires actively curating a blameless environment. He notes he still regularly meets people encountering blameless postmortems for the first time, thirteen or fourteen years after the concept entered the industry, including within Azure — which he points out is not a monolith but roughly 2,000 discrete services with different engineering teams, origin stories, and cultures.

No, the engineer was not fired. Klein was asked directly. Many highly impacted customers asked the same question during post-outage engagements, putting executives in the awkward position of saying "that is not what we do" without making the customer feel foolish for asking. Klein notes, wryly, that this is why he does not do those interviews.

The human toll is real and the system is responsible for it. Asked about the psychological weight on an engineer whose command caused billions of dollars of impact, Klein described his interview practice: he rarely talks to the engineer who performed the action first. He wants background and an understanding of the team's culture before going in. He has only encountered an engineer genuinely fearful of being fired once or twice, and either way he reinforces that firing is not what he is there for. His framing is unambiguous: if an engineer ran a command that took down the world, that is the system's fault for not protecting the engineer. The system failed the engineer.

The guardrail question has no clean answer. The same questioner raised the balance problem: there may be a handful of correct ways to execute a command and a thousand incorrect ones, so building code to block every wrong path means writing more prevention code than functional code. Klein does not offer a formula. His stated principle is that systems should help the engineer make good decisions and then protect the system when they do not, and that all systems should be designed for this — including giving the engineer insight into the current state of the system and the effect of the action they are about to take. That visibility framing is worth noting: it is a partial answer to the balance problem, because informing the operator is cheaper than enumerating every forbidden path.

Practical Takeaways

  • Watch for the narrative forming during the incident, not after. By the time the postmortem is written, "change gone wrong" or "operator error" may already be the accepted framing among people who were never on the bridge.
  • Treat every contributing factor as a repair candidate. Ask what had to be true or not true for the customer impact to occur in exactly the form it did, and enumerate those conditions rather than stopping at the last human action.
  • Use the diagram as an interview tool. Build it with one group, take it to another, and let them contradict it. Disagreement between teams surfaces the shared assumptions no single team can see.
  • Sort repairs by distance from impact. Near-impact nodes give you tactical fixes that prevent this specific recurrence; distant nodes give you the governance, tooling, and mental-model work that prevents a class of outages.
  • Audit your change-governance exemptions. The single most load-bearing factor here was the non-production classification, which silently disabled peer review, CAB review, emulation testing, and the listening period — for a device connected to the production IGP backplane.
  • Beware "non-production" that touches production control planes. Traffic isolation is not the same as control-plane isolation.
  • Sequence onboarding so guardrails land before operations. The AAA command audit existed as a process; it simply had not run yet on this role. Engineers reasonably assumed the guardrail was active.
  • In heterogeneous fleets, assume command semantics differ by vendor and OS version. Documented commands should record their blast radius per platform, not just their syntax.
  • Design SOPs for interpretation, not compliance. Tell operators what the command does and what risk it carries, so they have something to reason with when the instruction looks wrong.
  • Version and diff your SOPs, and surface changes at execution time. The engineer opened the SOP attached to the ticket with no signal that it had changed since they last used it.
  • Keep listening periods attached to risk, not to the production label.
  • Remember that automated recovery systems can amplify a total outage. Azure had to pause auto-reroute and auto-recovery during the event, and the pause itself extended the long tail.
  • Replace mandatory-training reflexes with targeted learning. Teaching a real incident during onboarding transfers the mental model; re-attesting to generic training does not.
  • Do not make culture change a precondition. Run the deeper analysis alongside the metrics reporting your leadership already expects.

Key Terms

  • Modern incident analysis — Klein's term for deep, systemic post-incident analysis of a selected class of incidents, contrasted with ITIL problem management and Five Whys derivatives.
  • DRI / OCE — Directly Responsible Individual or on-call engineer; the person paged when an alert fires.
  • SOP / MOP — Standard Operating Procedure and Method of Procedure; the written change instructions. Klein treats them interchangeably here.
  • Incident vs. outage (Azure usage) — An incident is a monitoring-threshold alert routed into a queue; an outage is a declared state with confirmed or imminent customer impact.
  • Sev 0 — An unofficial tier above Severity 1. It does not change the technical response but triggers additional communications and commitment workstreams.
  • CAB — Change Advisory Board; the review body in the SOP governance process.
  • Emulation environment — Microsoft's virtualization of the entire WAN inside Azure VMs, used to test SOP changes before production.
  • Contributing factor — An event or condition that had to be true, or not true, for the customer impact to occur in the form it did.
  • Counterfactual — A hindsight statement about what would have happened if a factor were removed; useful for generating repair candidates, not for establishing cause.
  • IGP (Interior Gateway Protocol) — The protocol family routers use to manage connectivity within a single network or autonomous system.
  • BGP (Border Gateway Protocol) — The protocol used to peer with external internet networks and exchange reachability information between autonomous systems.
  • Reconvergence — The process by which routers recompute the network topology and routing tables after a change or failure, during which forwarding can be unreliable.
  • Link state packet — An advertisement a router floods describing its own links, used by link-state IGPs to build the network map. Stale ones mean the map has diverged from reality.
  • AAA — Authentication, authorization, and accounting; the layer where Azure blocks commands identified as having more-than-local scope.
  • Re-IP — Changing the IP addressing of an already-provisioned device; routinely treated as a higher-risk operation by network engineers.
  • SWAN — Microsoft's Software-Defined Wide Area Network, the architecture that required these routers to move from external to internal IPs.
  • Listening period — The enforced observation window after a production change, during which key health indicators are watched before proceeding.
  • Blameless postmortem — The practice, popularized by John Allspaw's Etsy writing, of analysing incidents without attributing fault to individuals.

Where to Learn This

Asked what resources exist beyond the Five Whys practices most teams repeat, Klein pointed to the Resilience in Software Foundation community, the Adaptive Capacity Labs blog, the Learning From Incidents (LFI) community, and safety science books generally. He singles out the STELLA report, which he says he quotes from constantly — particularly a passage arguing that organizations prioritizing blame and accountability are deprioritizing the transfer of knowledge and the free flow of information. He also recommends the Howie Guide, originally from Jeli and now hosted by PagerDuty. The lineage, in his telling, begins with Allspaw's blameless postmortem writing at Etsy.

The through-line of the talk is that the January 2023 outage had exactly the shape the simple story described — an engineer ran a command and the world went dark — and that this description is nevertheless useless. Every fact in it is true and none of it tells you what to fix. The value of spending a month drawing an unreadable diagram is that it converts a story with one villain into a system with fifteen intervention points, most of which have nothing to do with the person who typed the command.


Reference: Sean Klein, The Time it Wasn't DNS, QCon San Francisco 2025, published by InfoQ on June 23, 2026.