Downtime is a period during which an IT system, server, application, website, network, or other service is fully or partially unavailable and cannot perform its intended functions. The metric is used to assess IT infrastructure availability and reliability, monitor SLA compliance, and analyze the impact of technical failures.
Downtime can be measured in seconds, minutes, or hours, or expressed as a proportion of the system’s total measured time. The shorter the total duration of downtime, the higher the actual service availability.
Types of Downtime
Depending on the cause and circumstances, several types of downtime can be distinguished:
- Planned downtime — a scheduled interruption for software updates, hardware replacement, maintenance, or other planned work.
- Unplanned downtime — unexpected unavailability caused by hardware failure, software errors, power outages, network problems, or other incidents.
- Full downtime — users completely lose access to the system or service.
- Partial downtime — the system continues to operate, but certain functions, components, or groups of users cannot access the required service.
When calculating SLA metrics, planned and unplanned downtime may be treated differently. The applicable rules are defined by the specific agreement.
Causes of Downtime
Downtime can result from both technical failures and errors in infrastructure operation. Common causes include:
- server, storage device, or network equipment failures
- power outages
- Internet connectivity and internal network issues
- software errors
- failed updates and configuration changes
- computing resource overload
- DDoS attacks and other cybersecurity incidents
- administrator and user errors
- data center incidents
- scheduled maintenance
The cause of an incident affects not only the duration of downtime but also the measures required to prevent similar incidents in the future.
How Downtime Is Calculated
Downtime can be calculated as the total duration of periods of unavailability within a selected time interval:
Downtime = Total Measured Time − Availability Time
If availability is expressed as a percentage, allowable downtime can be calculated using the following formula:
Downtime = Total Time × (1 − Availability / 100)
For example, with 99.9% availability over 30 days (720 hours):
720 × (1 − 0.999) = 0.72 hours, or 43 minutes 12 seconds.
For different availability levels over a 30-day period, the allowable downtime is approximately:
- 99% — 7 hours 12 minutes
- 99.9% — 43 minutes 12 seconds
- 99.99% — 4 minutes 19 seconds
- 99.999% — about 26 seconds
However, when calculating SLA compliance, the total time may be based not on the entire calendar period but only on the Agreed Service Time (AST).
Downtime and Uptime
Uptime represents the time a system is operational, while downtime represents the period when it is unavailable. If both metrics are calculated according to the same rules, together they make up the total measured time:
Uptime + Downtime = Total Measured Time
For example, if a service was available for 719 out of 720 hours in a month, its downtime was one hour and its uptime was 719 hours.
However, the technical uptime of hardware does not always reflect the availability of the end service. A server may continue to operate while an application hosted on it becomes unavailable due to a software error. From the users’ perspective, this is application downtime even though the server itself has not stopped running.
Downtime and SLA
An SLA (Service Level Agreement) typically defines the acceptable service availability level and the rules for measuring downtime. Therefore, assessing SLA compliance requires considering not only the actual duration of downtime but also the conditions under which it is counted.
The agreement may specify:
- the minimum guaranteed availability level
- how the start and end of an incident are recorded
- the period over which the metric is calculated
- which types of downtime are included or excluded
- rules for scheduled maintenance
- compensation or service credits for SLA violations
For example, a provider may exclude pre-agreed maintenance windows from the calculation. In this case, the actual period of service unavailability and the downtime counted for SLA purposes will differ.