Operations Guide

Running a data centre: operations, resilience & uptime

Commissioning a facility is the start line, not the finish. Uptime isn’t something you buy with redundant kit — it’s something you earn, every day, through disciplined operations. This guide covers the availability stack, what the “nines” actually cost you in downtime, the operational disciplines that separate a resilient site from a fragile one, and how incidents are really handled.

A vendor-neutral, educational guide. Availability figures and frameworks are referenced from the Uptime Institute, ASHRAE and common industry practice; see notes at the end.
≈66%
Of significant outages involve human error — almost always a skipped or absent procedure, not bad luck
1.6hrs
Annual downtime at 99.982% — the availability most enterprises target
5min
All the downtime “five nines” (99.999%) allows you in a whole year
MTTR
The one availability lever operations owns outright — how fast you recover, not how much plant sits underneath
24⁄7
The operational reality: monitoring, response and on-call never stop, including the 3am of a bank holiday
01 | The Reality

Uptime Is Operated, Not Bought

The building gives you power, cooling and space — but that’s the floor, not the finish. Whether your estate actually stays available depends on how it’s operated, day in, day out. And the industry’s own data is blunt about it: most major outages aren’t plant failures — they’re operational. A change made without controls, an alarm no one answered, a dependency nobody documented. That operational layer — not the M&E underneath it — is where availability is won or lost, and it’s the layer this guide is about.

The facility is the floor

Power, cooling and space are the facility’s M&E — they set the conditions, but they don’t keep your estate running. Everything above the floor — hardware, connectivity, changes, monitoring — is operations, and it’s where the day-to-day risk lives.

People are the variable

The single largest cause of avoidable downtime is human error during change and routine work. Procedures (MOP/SOP/EOP), training and supervision are the controls that turn a good environment into a reliable one.

The risk is in the gaps

Outages cluster around the seams: an undocumented change, an un-diverse cross-connect, a monitoring blind spot, a single engineer who knows how the rack is wired. Operations is the discipline of closing those gaps before they find you.

Recovery is a capability

Failures are inevitable; extended outages are not. Whether a fault is a 30-second blip or a six-hour incident is decided by monitoring, escalation and rehearsed response — long before the alarm sounds.

02 | The Stack

The Operational Stack

The facility hands you power, cooling and space. Operations is everything done on top of that to keep your estate available — the hardware, the connectivity, the capacity, the monitoring and the people. The estate is only as available as the weakest of these, so each is a domain an operations team runs around the clock.

A

Hardware & rack operations

Installs, moves, swaps, break-fix and RMA handling. The hands-on work of bringing kit in, keeping live hardware healthy, and replacing what fails — correctly and to schedule.

B

Connectivity & cabling

Cross-connects, patching, port management, labelling and documentation. Power keeps kit alive; the network makes it useful — and undisciplined cabling is a slow-motion outage waiting to happen.

C

Capacity & asset management

Tracking space, power and port headroom against an accurate, live asset register — so you never strand capacity or, worse, run out of it mid-deployment.

D

Monitoring & alerting

Watching the environmental and power telemetry the facility exposes — temperature, humidity, load, access — and acting on it. Operations spots the deviation and escalates while there’s still time.

E

Physical security & access

Access control, escorting, CCTV review and audit logging. Who is on the floor and what they touch is part of availability and of compliance — and a discipline in its own right.

!

Single points of failure

The job is to find the link that isn’t resilient — the un-diverse cross-connect, the undocumented dependency, the single competent engineer. Unmapped SPOFs are where “resilient” estates still fail.

03 | Self-Assessment

How Operationally Mature Are You?

Operational maturity is the gap between a site that has good kit and one that is run well. Score yourself honestly against the disciplines below — the bars show where most outages are won or lost. Anything you can’t evidence with a document or a record is a gap, not a strength.

Monitoring & alerting
Critical
Procedures (MOP/SOP/EOP)
Critical
Change control
Critical
Connectivity & cabling
High
Incident response & on-call
High
Capacity management
High
DCIM / asset accuracy
High
Failover & DR testing
Medium
Training & competency
Rising

Monitoring and procedures sit at the top deliberately — you can’t respond to what you can’t see, and you can’t act safely without a method. The most common failure pattern isn’t missing redundancy; it’s redundancy that was never tested, or a procedure that existed only in one engineer’s head.

04 | The Numbers

What the “Nines” Really Cost

Availability is quoted in “nines,” and the difference between them is brutal: each extra nine cuts your allowable downtime roughly tenfold. Knowing what a target actually permits — in minutes per year — is the only honest way to decide what resilience you need and what it’s worth paying for.

AvailabilityDowntime / yearDowntime / monthTypical context
99% (“two nines”)3.65 days7.2 hrsNon-critical / dev & test
99.9% (“three nines”)8.76 hrs43.8 minStandard business systems
99.982%1.6 hrs~8 minCommon enterprise availability target
99.99% (“four nines”)52.6 min4.4 minCritical production
99.999% (“five nines”)5.26 min26 secMission-critical / fault tolerant

These are the maths of availability, not a guarantee — a target only holds if the operations behind it do. “Five nines” means roughly five minutes of unplanned downtime in an entire year, which takes a resilient facility and, just as importantly, mature operations — and operations is the part most under-invested in. Note too the distinction between planned and unplanned downtime: a well-run estate schedules disruptive work into change-controlled maintenance windows, so the goal becomes zero unplanned outage.

MTBF — how often

Mean Time Between Failures: the expected interval between faults on a component. Higher is better, and it’s improved by quality kit and preventive maintenance.

MTTR — how fast

Mean Time To Repair/Recover: how long to restore service once something fails. This is the number operations controls most directly — through spares, monitoring and rehearsed response.

Planned ≠ unplanned

Disruptive work scheduled into a change-controlled maintenance window is planned and managed. Counting that as “downtime” conflates it with the real enemy — the goal is zero unplanned outage.

05 | The Disciplines

The Operational Disciplines

Reliable sites aren’t reliable by accident. They run a set of repeatable disciplines that catch problems early, make every action safe and repeatable, and keep the facility’s capacity ahead of demand. These are the practices that turn a redundant design into a resilient operation.

Change control

No change to a live system without review, approval, a method statement and a rollback plan. The single highest-leverage control against self-inflicted outages — because most outages are self-inflicted.

MOP / SOP / EOP

Method, Standard and Emergency Operating Procedures: the written, rehearsed scripts for routine work and for the bad day. When the alarm goes at 3am, you follow the EOP — you don’t improvise.

Capacity management

Tracking space, power and port headroom so you never strand capacity or, worse, run out of it mid-deployment. Stranded capacity is wasted spend; exhausted capacity is a stalled project.

Asset lifecycle management

Knowing what every asset is, where it sits and where it is in its life — from receipt and deployment through refresh to secure decommission. The backbone of capacity, support and audit.

Connectivity & cabling discipline

Structured cabling, change-managed patching, labelling and accurate records. Tidy, documented connectivity is fast to work on and safe to change; the alternative is an outage nobody can trace.

DCIM & documentation

An accurate, live record of every asset, circuit and connection. Decisions made on stale data cause outages; DCIM and an up-to-date asset register are the operational source of truth.

When something goes wrong: the incident lifecycle

Detection-to-recovery is a rehearsed sequence, not a scramble. Every minute saved here is a minute off your MTTR — and the difference between a near-miss and a reportable outage.

Detect

Monitoring catches the deviation — a temperature trend, a failed unit, a power anomaly — ideally before it affects the load.

Alert & triage

The alarm reaches the right person on a clear escalation path, and severity is assessed against defined thresholds.

Respond to procedure

The relevant EOP/SOP is executed — isolate, fail over, switch to alternate path — methodically, not improvised.

Stabilise

Service is protected or restored, the fault is contained, and the site is returned to a known-good, redundant state.

Restore & verify

The failed component is repaired or replaced, redundancy is re-established, and the fix is verified under load.

Review

A blameless post-incident review feeds back into procedures, training and design — so the same fault can’t bite twice.

06 | The Bridge

How Optronix Helps

Operations is the longest stage of the data centre lifecycle — the years between building a facility and decommissioning it — and the one where reputation is made or lost. We support clients across the operational disciplines, with smart hands and live-operations cover, across the UK, Europe and worldwide — whether you run the site yourself or want us carrying the load.

Operational readiness & audit

We assess your estate against the disciplines above — procedures, change control, monitoring, capacity, cabling and SPOFs — and give you a prioritised plan to close the gaps.

Smart hands & remote ops

Eyes and hands on the floor for the lights-out site: installs, swaps, audits and break-fix, to your procedures and to schedule.

24/7 monitoring & response

Round-the-clock cover with clear escalation and rehearsed incident response — so an alarm at 3am is met with a procedure, not a scramble.

Capacity, asset & cabling support

Capacity tracking, asset-register and DCIM accuracy, and structured-cabling work — keeping your estate ahead of demand and audit-ready.

Running a facility — or stretched thin running it?

Tell us where it hurts — coverage gaps, an upcoming maintenance window, a site that needs eyes on it — and we’ll put the right operational support around it, on the ground wherever it is.

Talk to our team

Sources & notes

  • Uptime Institute Annual Outage Analysis — on human error as a factor in the majority of significant outages, and the planned/unplanned distinction
  • Availability / downtime figures are the standard arithmetic of the “nines” (e.g. 99.999% ≈ 5.26 min/year)
  • Common industry practice for change control, MOP/SOP/EOP, DCIM, and capacity & asset management