System status

Checking every system…

Live from the probe record. This page reads the same database the alerts fire from — there is no second, prettier copy of these numbers. Nodes are the machines your box runs on; website & API is how you reach us. They are reported apart because they fail apart.

Customer nodes · 90 days
Website & API · 90 days
Since the last bad minute
Last probe
Ninety days, per system

What we measure

One bar per day. Green is a day with no failed probe; amber is a wobble; red is real downtime. Hollow bars are days we have no record for — they are never quietly painted green.

Loading the probe record…
Operational Blip < 1 min Degraded Outage No data Reconstructed from system journals
Awkward comparison

90-day uptime, their number

The figure each status page publishes about itself — the "99.41% uptime" printed under its own 90-day chart. It is the only peer number directly comparable to ours: same measurement, same window. We show the worst component on each page, and open any row to see every component behind it.

We rank on this rather than on declared incidents, because counting announcements measures how talkative a status page is, not how much it was up — Twilio declares an incident for every regional wobble and tops that table while publishing 99.97%. Roughly half of these pages switch the showcase off; those rows say "not published" rather than having a number invented for them. Declared incidents and our own probe measurements are still in every row.

Fewer days is better. Board rotates to whoever has been worst.
Reading eighteen status pages…

Nothing deleted

Incident log

Everything the record knows about, worst days included. A status page you can only ever screenshot when it is green is an advertisement, not a status page.

Loading…
How this is measured

The boring part

Read this before you believe any of the numbers above — including ours.

Who is doing the measuring?+
Two probers, one minute apart, on two different networks in two different cities — one on the controller in New York, one off-site in Los Angeles on an unrelated provider. A system is only recorded as down when both agree it is down, so one prober losing its own uplink never becomes our outage. The off-site prober keeps writing to disk when it cannot reach us and uploads the backlog afterwards, which is exactly how the minutes where we were unreachable end up on this chart.
Does a missing probe count as uptime?+
No. That is the standard way a status page lies. A minute with no probe result is excluded from the maths and drawn as a hollow bar, so the percentage is always "of the time we actually watched", never "of the time we felt like counting".
What are the reconstructed days?+
The probe fleet is newer than the company. For the days before it existed we rebuilt downtime from the machines' own system journals — boot records and service stop/start transitions — which is evidence, not estimation. Those days are marked separately, and reconstruction can only ever add known-bad time: it has no path to turn an unknown day green.
Is this comparison actually fair?+
It is handicapped three ways in their favour. Their planned maintenance is excluded. The window is capped to what every feed still covers, so nobody gets credit for incidents that scrolled off the end. And their incidents are often scoped to one region or one product, while ours count against the whole company. We are also, obviously, a much smaller target — a rack and a fibre line versus a global network. Fewer moving parts is a real advantage and we will take it, but it is the reason, and pretending otherwise would be the sort of thing this page exists to avoid.
What happens when you do go down?+
It appears here, in red, without being asked. Boxes are backed up nightly to off-site storage, the controller has an automatic standby in another datacenter, and there is a human on Discord rather than a ticket bot. See Service Levels for what we actually commit to in writing.

Up when it matters.

That is the entire product. Pick a box and watch it go live.