The First Hour When Your Site Goes Down
It almost never starts with an alert. It starts with a customer calling to say she tried to place an order last night and got an error page, or a supplier mentioning that his email to you bounced back. That call is usually the first you hear when your site goes down, and the hour that follows will go one of two ways depending on what you do with the first four minutes of it.
The instinct is to start fixing, and it is wrong for four minutes
The urge when the phone rings is to do something. Restart the thing you know how to restart. Log in and click around. Text the person who built the site. Send the host a message that says site down please help. All of that feels like urgency, and none of it is diagnosis.
Two businesses can go down for the same reason on the same Tuesday, and one is back in twenty minutes while the other loses two days. The difference is rarely skill or money. It is the order they did things in. The one that lost two days started repairing before it knew what was broken, changed nameservers on a hunch, restored a backup over data that was fine, and ended up with two problems instead of one.
Four minutes of triage is not slow. It is the only part of the hour that is actually under your control. Server hardware, a registrar's queue and a mail provider's spam desk all move at their own speed no matter how loudly you ask. What you decide in those four minutes points everything else in a direction.
Repairing before you know what is broken is how one outage becomes two.
Down for you is not the same as down
Start with the least flattering possibility, which is that nothing is wrong. The customer who called is sitting on one network, with one DNS resolver, one browser cache and possibly one office firewall appliance that decided your certificate looked suspicious. She is telling the truth about her experience. That is not the same as telling you about your server.
The test takes thirty seconds. Pick up a phone, turn wifi off so it is running on cellular data, and load the site. Different network, different resolver, different cache. If it loads on cellular and fails on the office wifi, the site is up and the problem lives inside that building, which makes it a call to their IT person or their internet provider and not to us.
This matters more than it sounds like it should, because the answer decides who you wake up. Say the unwelcome part plainly to the person who called: it is loading here and on three other networks, so let us look at yours. People take that fine in the first five minutes. They take it badly after an hour of watching you pretend to fix something.
The most useful minutes when your site goes down are the boring ones
Once you know it is down for everybody, there is one question worth answering before any repair starts. Your website, your domain and your business email feel like a single thing because they all live behind the same name. They are three services and they fail separately.
Load the site and note exactly what happens, because a page that will not connect, a 500 error, a browser warning about an expired certificate and a registrar parking page are four different outages with four different owners. Then send yourself a message from an outside address, a personal Gmail account is fine, and see whether it lands. Then check whether the domain answers at all.
That is the whole triage. It is boring, it takes about four minutes, and it decides everything after it. If the website and the email went dark inside the same minute, stop looking at the website. Two services that share nothing but a domain name do not fail together by coincidence.
The triage is boring, which is exactly why people skip it.
Three outages, three different clocks
Knowing which one you have tells you how long you will be down, which is what your customers and your staff actually want from you. A server problem is usually the shortest clock. Something gets restarted, or a full disk gets cleared, and you are counting in minutes.
DNS runs on a slower clock, and the delay is designed in. Every DNS record carries a time to live, often an hour, sometimes a day. If a record was wrong and you correct it, anyone whose resolver already cached the wrong answer keeps getting the wrong answer until that timer runs out. Nothing is broken during the wait and there is nothing to do about it, which is a hard thing to explain to somebody staring at a blank page.
A lapsed domain is the worst clock, because the problem is not on your server at all, it is a state at the registry. Renewing inside the grace period usually puts things back quickly. Past that it moves to redemption, which costs a fee well beyond the renewal price and takes days, and the site plus every mailbox on the domain stays dark the whole time. Read what happens when a domain expires on a calm afternoon, because once that timeline starts it is not negotiable.
An expired certificate is the friendliest of the four. The server is fine, the site is up, and every visitor gets a full red warning telling them not to trust you. Ten minutes of work, and the ugliest looking hour on the list.
Who you call depends on what you found
Site down, please help is the least useful message you can send anybody. It contains no fact that narrows anything, so the person receiving it spends the first ten minutes asking you the questions you could have answered before you hit send.
What moves things along is the exact address you tried, whether it fails on more than one network, the exact error text or a screenshot of it, the time it started, and what changed the day before. A registrar parking page is not something a hosting technician can fix, and a 500 error coming out of your own site is not something a registrar can fix, and either one aimed at the wrong desk buys you an afternoon of being politely transferred. Bring those five facts to hosting and domain support and the first reply is usually an answer instead of another question.
That last item, what changed, is the one people skip and the one that solves the most outages. Very few sites go down on their own. A plugin updated itself at 3am. A developer pushed a change at 4:45 on a Friday. A card on file was reissued after a fraud alert, the automatic renewal failed quietly in April, and the notice went to an address belonging to somebody who left the company in 2023. Outages have authors, and the author is usually a calendar.
The bill for the hour arrives later
Uptime numbers make this easy to underrate. 99.9 percent of a year is 8.76 hours, so one bad afternoon spends the whole annual allowance. 99.9 percent of a thirty day month is about 43 minutes. You do not have much outage time to work with, which means the ratio of minutes spent diagnosing to minutes spent repairing is most of what sets the final number.
The website is also the cheap half. An hour of downtime on a Tuesday afternoon costs some traffic and some irritation, and most people who hit a dead page try again later. Mail is different. When mail is down there is no error page for anybody to see. The person who wrote asking for a quote gets nothing back and quietly decides you could not be bothered to answer, and you never find out it happened. That is why mail belongs in the first four minutes and not in the second hour.
There is no error page for email, so the sender just decides you did not answer.
Nobody is graded on how fast they started. They are graded on when the site came back, and on whether anything else got broken along the way. Four minutes spent establishing whether it is down for everyone, and whether the failure is the site, the domain or the mail, is the cheapest four minutes you will spend all year. Everything after those four minutes is either the right work or the wrong work, and by then it has already been decided.
More from the blog
- What Unlimited Hosting Really Means
- Why Your Business Email Goes to Spam
- What a Traffic Spike Actually Does
All posts are on the blog. For the step by step versions, see the domain and hosting guides.