When Your Server or System Goes Down: Emergency Triage Steps
Emergency triage for sudden server or business-system outages: narrowing causes with 'one person or everyone?', safe restart steps and when not to restart, who to contact with or without a maintenance contract, and post-recovery actions.
The Order of Diagnosis (Bottom Line)
When a server or business system suddenly goes down, don't start poking at hardware randomly. Checking things in the following order is the fastest path to a diagnosis. Restarting several devices at once or repeatedly unplugging and reconnecting cables out of panic will only make the cause harder to find and delay recovery.
- ① Check the scope of impact (is everyone affected, or just some people?) — getting this wrong often leads you to suspect the wrong thing
- ② Check the physical state of power and network connections — a loose plug, a bad cable connection, or a tripped breaker are common, simple causes
- ③ Check the state of the hardware (the server itself, routers, etc.) — look at status lights and any error indicators
- ④ Check for problems on the software/application side — if the hardware looks fine, suspect an application or database error
Narrow Down the Cause by Who and Where Is Affected
The fastest route to a diagnosis starts with the scope of impact. Simply asking "is everyone affected or just one person?" and "is it just the internal system, or is the internet connection itself down?" narrows down where to look considerably. Use the table below to find the pattern closest to your situation.
| Symptom Pattern | Likely Cause | Next Step |
|---|---|---|
| Nobody can connect, internal network is also down | Line outage or power equipment issue | Contact your line provider, check the breaker panel |
| Nobody can connect, internal network is fine | Server hardware or system-side failure | Check server status, contact your maintenance vendor |
| Only one location is affected | Router or line issue at that location | Restart equipment at that location, contact the line provider |
| Only one person (one device) is affected | PC hardware or account issue | Restart that PC; for details see diagnosing PC/network problems |
| Only one specific business system is down | Application or database failure | Check error logs, contact your system vendor |
How to Restart Safely
If your diagnosis points to a problem with the hardware itself and there's no sign of unauthorized access, restarting is often the simplest fix. But pressing the power button without following the steps below can make later investigation much harder.
- First, record the error message and screen state with a photo or notes (an important clue for later investigation)
- Notify affected users of the temporary disruption (giving them time to save any work in progress)
- Restart the server using its normal shutdown procedure (a forced power cut should be a last resort — it raises the risk of data corruption)
- After restarting, verify the integrity of any data that was being processed (check that order or accounting data wasn't left mid-transaction)
Cases Where You Should Not Restart
Even if the problem looks like a simple glitch, do not restart if any of the following apply — contact your maintenance vendor or an expert first. A hasty restart can make the damage worse or make the root cause impossible to determine.
- Unauthorized access such as ransomware is suspected (see emergency response to ransomware for details)
- An error occurred while data was being written to a database
- The system keeps crashing for an unknown reason (possible hardware failure — repeated restarts can cause further damage)
- A process is running whose data could be lost if the system restarts
- There are signs of physical trouble, like a burning smell or unusual noise (don't force the power on — contact the manufacturer instead)
Contacts Depend on Whether You Have a Maintenance Contract
If you have a maintenance contract, contact the desk listed in the agreement and check whether an after-hours emergency contact exists, along with the expected response time (SLA). If you don't have a contract, or can't reach anyone, refer to what to do when you can't reach your vendor to find an alternative contact. If nobody on staff is deeply familiar with the system, it's also important to decide early not to force a DIY fix and instead bring in outside help sooner rather than later.
What to Do After Recovery
- Record the cause and the actions taken in chronological order (useful reference if the same symptom appears again)
- Confirm that backups are being taken correctly (the outage itself may have interrupted your backup process)
- Consider configuration changes or improved monitoring to prevent the same issue from recurring
- If partners or customers were affected, once things have settled, briefly explain what happened and what you're doing to prevent it recurring
Preventing Recurrence
To avoid repeating ad-hoc recoveries, it's essential to review your ongoing maintenance setup. With regular inspection and monitoring in place, many failures can be caught as early warning signs before a system "suddenly" goes down. See the complete guide to IT maintenance and operations for an overview of how to structure it. For initial response to other kinds of IT trouble — such as email outages or website issues — see the IT emergency first response guide as well.
Will turning the power off and on again fix it?
That can work for a simple freeze, but repeating it without identifying the cause can make things worse. Always record the system's state before restarting, and suspect a hardware failure if the problem keeps recurring.
We don't have a maintenance contract and don't know who to contact.
Start by contacting the vendor or system company that set up your system. If they don't respond, see what to do when you can't reach your vendor for alternative contacts.
If the outage drags on, how should we explain it to business partners?
Promptly share the scope of impact, current response status, and expected recovery time to the extent you know it — this maintains trust. Avoid speculating about the cause; share only confirmed facts.
Even if we manage to recover the system ourselves, should we still consult an expert?
Even after an emergency fix, leaving the root cause unidentified leaves a risk of recurrence. It's worth having an expert investigate the underlying cause even after you've recovered.
Related free tools (no sign-up, instant results)
Feel free to contact us
Contact Us