Skip to content
Blog

Outage report #1: what 225,000 daily probes actually see

We audited our own outage classifier and found 90.5% of our raw down readings were our probes failing, not vendors. Here is what we found, what we fixed, and what actually happened.

We run public probes against 940 services in our Outages catalog, from ten regions, around the clock. On 2026-08-27, when we ran the audit behind this report, the catalog held 195 services and we were probing from four of those regions, logging 225,421 catalog probes in the preceding 24 hours. We also poll vendor status pages directly and try to read what they say: over that same 24 hours, 25,046 polls, 15,746 of which returned a claim we could actually parse into our truth scorecard.

That is a lot of data, and for a while we trusted it more than we should have. This is the report on what we found when we stopped trusting it and checked.

What fooled us

We run an internal audit of our own outage classifications, the kind we'd want a competitor to run on us. On 2026-08-27 we ran it for the first time end to end and pulled 2,117 raw "down" probe readings for review.

90.5% of them were not vendor outages. They were us.

  • 47.0% were our own 10-second timeout expiring with no HTTP response at all. A slow or blocking network path in front of a vendor looks identical to a real outage if your only signal is "did a response arrive."
  • 39.7% were redirect loops, almost all of them aimed at datacenter IP ranges. A chain of redirects that never terminates on a normal host is a strong sign our probe was being routed away from the real endpoint, not that the endpoint was down.
  • 3.8% were wrong-endpoint 404s and 406s, wrong-endpoint errors, not outage errors.

Only 4.2% of the raw readings were genuine 5xx server errors from the vendor itself.

The pattern behind almost all of it is bot-blocking. Datacenter IP ranges get challenged, redirected, or timed out by design, by the exact same infrastructure that protects a site from scrapers. An automated prober sitting in a cloud region looks a lot like a bot, because it is one. Every "is it down" checker that runs from datacenter IPs is measuring some mixture of "is the vendor down" and "did this vendor's bot-defense notice me today," and most of them don't say which.

What we changed

We rewrote the classifier rather than patch around the symptom:

  • HTTP 406 and 451 responses are now recorded as refusals, not outages. A server that is up and deliberately declining to answer a request is not the same event as a server that has fallen over.
  • Timeouts that correlate across multiple unrelated services in the same probe window are no longer counted as an individual service being down. A correlated hang is evidence about our own network path, not about that one vendor.
  • Readings shorter than our own timing resolution are labeled as such instead of being forced into a definite up or down bucket.

We then applied the rewritten classifier to history, not just to new data. That pass reclassified 269 raw readings and retracted 25 published outage events that no longer met the bar. Those retractions are the point of publishing this: a monitoring product whose worst content is confident about mistakes it never corrects is exactly the kind of "is it down" data everyone learns to distrust. We would rather show the correction than hide the original error.

What actually happened on the internet, in the two weeks before the audit

After the reclassification, exactly six events in the catalog survive from the week of 2026-08-17 through 2026-08-26:

  • Week of 08-17: four events
  • Week of 08-24: two events

Five of the six are Paramount+ down windows in our iad (Virginia) region only, 30 to 62 minutes each. The sixth is Microsoft 365, in our fra (Frankfurt) region, 10 minutes, on 2026-08-26.

A caveat we want on the record, not buried: the Paramount+ events and the Microsoft 365 event are single-region readings from our own probes, not vendor-confirmed outages. We saw a service stop responding correctly from one vantage point for a bounded window. We did not get, and are not claiming, independent confirmation from the vendor or from a second region. Present these as what our probes measured from where they sit, not as confirmed outages, because that is what the data actually supports.

Method notes

  • Coverage at the time of this audit (2026-08-27): 195 active services in the public catalog. The catalog has since grown past 900; see the 2026-08-31 post on the curation behind that.
  • Cadence at the time of this audit (2026-08-27): continuous probing from four regions -- iad, sjc, fra, nrt -- producing 225,421 catalog probes over the trailing 24 hours.
  • Cadence today: the fleet has since expanded to ten regions -- iad, sjc, fra, nrt, ord, yyz, lhr, sin, syd, gru -- the count live on realuptime.io today. We have not re-run this audit against the ten-region fleet; the classification numbers above are the four-region measurement, unchanged.
  • Vendor status-page polling: 25,046 polls in that same 24-hour window (2026-08-27), 15,746 of which returned text we could parse into a structured claim for the truth scorecard that compares what vendors say against what we measure.
  • What "blocked" means here: a probe response that indicates the request was intercepted, challenged, or redirected by bot-defense infrastructure rather than answered by the service itself -- the timeout, redirect-loop, and wrong-endpoint categories above.
  • Single-region events: flagged explicitly wherever they appear, because a one-vantage-point reading is a different, weaker claim than a multi-region confirmation, and we don't think the difference should be flattened away in the copy.

This is the first in a periodic series. We'll keep publishing it on the same terms: real probe counts, real classification errors when we find them, and events labeled by exactly how much confirmation they actually have.