Original research

The state of incident response

We read the published incident history of 77 software providers — 2,544 incidents over twelve months — to find out how long outages actually run, and what people are told while they wait.

The medians look reassuring. Almost nothing interesting is in the medians.

15h 01m
of extra open time the "minor" label buys, in the bad case
12.9x
longer the worst waits between updates are than the typical one
36%
of incidents ended with no cause given at all
4.8x
more incidents start on a weekday than at the weekend

What the word “minor” costs you

Critical incidents are resolved faster than minor ones. That is not a finding, it is what triage is for, and the typical case barely shows it: a median of 1h 25m against 2h 02m is close enough to be noise.

The gap opens in the tail. At the 90th percentile a critical incident has been running for 8h 06m, and a minor one for 23h 07m — 2.9x longer, and 15h 01m of additional open time. A tenth of the incidents labelled minor were still open the better part of a day later.

That is a statement about attention rather than difficulty. “Minor” is the label an incident gets when nobody is going to be woken up for it, and the tail is where that decision shows up.

SeverityIncidentsMedian90th percentile
critical1521h 25m8h 06m
major5051h 49m12h 20m
minor1,1652h 02m23h 07m
Unclassified4951h 01m13h

Severity is the provider’s own label, not our reclassification. “Unclassified” is the API’s none impact, used for informational posts and for incidents a provider never graded.

The silence is worse than the outage

Half of all updates land within 46m of the previous one, which reads like an industry doing a decent job of keeping people informed. Then the tail: 9h 52m at the 90th percentile. 12.9x the typical wait.

That is the widest spread in the whole dataset — wider than time-to-resolve, which runs 1h 46m at the median and 18h 02m at p90, a spread of 10.2x. Communication degrades faster than the incidents themselves do. The people worst affected by an outage are the ones being told least about it, because a long incident is also a busy one and status updates are the first thing to be dropped.

The median incident got 3 updates. 8% got exactly one and no follow-up: posted, then nothing until it was marked resolved.

A third of the time, nobody says what broke

In 36% of incidents — 917 of 2,544 — the provider never said what had gone wrong. No identified status, and no explanation anywhere in the update text either.

This figure is smaller than the one we first calculated, and the difference matters. Counting only incidents that never reached the identified status gives 54%, which is a much better headline and a worse number: around a third of those explained the cause in prose without ever moving the status. Counting them as silent would have measured a workflow convention and called it secrecy. 36% is what survives checking.

Where a cause was identified it came quickly — a median of 17m — but the tail runs to 2h 12m. Mitigation followed at a median of 54m.

Outages follow change, not traffic

Incidents begin on weekdays at 4.8x the weekend rate — 469.4 a day against 98.5. Weekend traffic does not fall by anything like that much for most of these providers, so this is not load. It is deployment, migration and configuration: the things people do during the week.

Day started (UTC)Incidents
Mon476
Tue475
Wed457
Thu527
Fri412
Sat107
Sun90

How long incidents actually run

Of the 2,317 incidents with both a start and a resolution timestamp:

Time to resolveIncidentsShare
Under 30 minutes51522%
30 to 60 minutes30613%
1 to 4 hours83736%
4 to 12 hours35815%
Over 12 hours30113%

Roughly a third are dealt with inside the hour. The last 13% run past twelve hours, and those are where the communication tail above lives.

Method

Every provider in the sample publishes a machine-readable incident history through the Statuspage v2 API, at {host}/api/v2/incidents.json. We read that endpoint once per provider, kept every incident that started between 2025-09-07 and 2026-09-07, and discarded planned maintenance. That leaves 2,544 incidents across 77 providers.

Durations are measured from the incident’s own created_at to its resolved_at. “Time to identify” uses the first update whose status is identified, and “mitigation” uses monitoring_at. Gaps between updates are measured between consecutive posts on the same incident.

None of this uses NoDisrupt customer data. That is deliberate twice over: a number a reader cannot re-derive is worth little, and publishing statistics drawn from customers’ own monitoring would put their operational detail into a public document.

What this data cannot tell you

There is no incident-frequency figure here, and there should not be. The API returns only the 50 most recent incidents per provider. 33 of the 77 providers hit that cap inside the window, so for those we are seeing a truncated slice rather than a complete history. Averaging a rate across truncated and untruncated providers would understate precisely the busiest ones. Durations and communication behaviour are per-incident measurements, so the cap does not affect them.

The sample is not the industry. It is providers using one hosted status-page product. Companies on other status vendors, or on bespoke pages, are absent — and so is every company that publishes no incident history at all, which is most of them. If anything that biases the sample toward the transparent end of the market, which makes the 36% with no stated cause more surprising rather than less.

Severity is self-reported. One provider’s major is another’s minor, and the incentive runs one way. We did not reclassify anything, so the severity table measures how providers label and respond to their own incidents, not an objective scale of harm.

The 77 providers in the sample

Listed so any figure above can be checked against the source.

  • bitbucket.status.atlassian.com
  • status.1password.com
  • status.airtable.com
  • status.amplitude.com
  • status.appcenter.ms
  • status.asana.com
  • status.bitrise.io
  • status.box.com
  • status.braze.com
  • status.browserstack.com
  • status.bugsnag.com
  • status.buildkite.com
  • status.circleci.com
  • status.cloudbees.com
  • status.cloudinary.com
  • status.coinbase.com
  • status.confluent.cloud
  • status.contentful.com
  • status.crates.io
  • status.customer.io
  • status.datadoghq.com
  • status.digitalocean.com
  • status.discord.com
  • status.doppler.com
  • status.dropbox.com
  • status.duo.com
  • status.elastic.co
  • status.epicgames.com
  • status.expo.dev
  • status.figma.com
  • status.fly.io
  • status.gocardless.com
  • status.grafana.com
  • status.hashicorp.com
  • status.hubspot.com
  • status.jfrog.io
  • status.kinsta.com
  • status.klaviyo.com
  • status.launchdarkly.com
  • status.linode.com
  • status.logdna.com
  • status.mailgun.com
  • status.miro.com
  • status.mixpanel.com
  • status.mongodb.com
  • status.netlify.com
  • status.newrelic.com
  • status.npmjs.org
  • status.onesignal.com
  • status.opsgenie.com
  • status.optimizely.com
  • status.percy.io
  • status.plaid.com
  • status.prismic.io
  • status.pulumi.com
  • status.pusher.com
  • status.render.com
  • status.rollbar.com
  • status.rubygems.org
  • status.sanity.io
  • status.saucelabs.com
  • status.segment.com
  • status.sendgrid.com
  • status.sentry.io
  • status.snowflake.com
  • status.snyk.io
  • status.sumologic.com
  • status.supabase.com
  • status.testrail.com
  • status.twilio.com
  • status.twitch.tv
  • status.vercel.com
  • status.wise.com
  • status.zapier.com
  • status.zoom.us
  • www.cloudflarestatus.com
  • www.githubstatus.com

Questions

Where does this data come from?

The public incident history of 77 hosted status pages, read from the Statuspage v2 API at {host}/api/v2/incidents.json. It is public data on a documented endpoint, so any figure here can be re-derived independently. None of it comes from NoDisrupt customers.

Why is there no 'incidents per month' figure?

Because it would be wrong. The API returns only the 50 most recent incidents per provider, and 33 of the 77 providers hit that cap inside the twelve-month window, meaning their history is cut off rather than complete. Averaging a rate across truncated and untruncated providers understates exactly the busiest ones. Durations and communication behaviour are per-incident measures and are unaffected.

Does this mean these providers are unreliable?

No, and it is close to the opposite. Every provider here publishes a status page with a machine-readable history, which is more transparency than most software companies offer. The sample is drawn from organisations that chose to be measurable.

What counts as an incident?

Anything the provider itself posted to its status page and later resolved, at any severity. Planned maintenance is excluded throughout. We take the provider's own severity label rather than reclassifying anything.

Why we looked

The tail is a monitoring problem

Nothing in this data is about how hard these outages were to fix. The long incidents are the ones that were noticed late, diagnosed slowly, or quietly deprioritised — and each of those is decided before any repair work starts. That is the part monitoring and escalation actually move.

For the arithmetic behind an uptime commitment, see the uptime calculator and the uptime monitoring guide.