Running a data centre: operations, resilience & uptime
Commissioning a facility is the start line, not the finish. Uptime isn’t something you buy with redundant kit — it’s something you earn, every day, through disciplined operations. This guide covers the availability stack, what the “nines” actually cost you in downtime, the operational disciplines that separate a resilient site from a fragile one, and how incidents are really handled.
Uptime Is Operated, Not Bought
The building gives you power, cooling and space — but that’s the floor, not the finish. Whether your estate actually stays available depends on how it’s operated, day in, day out. And the industry’s own data is blunt about it: most major outages aren’t plant failures — they’re operational. A change made without controls, an alarm no one answered, a dependency nobody documented. That operational layer — not the M&E underneath it — is where availability is won or lost, and it’s the layer this guide is about.
The facility is the floor
Power, cooling and space are the facility’s M&E — they set the conditions, but they don’t keep your estate running. Everything above the floor — hardware, connectivity, changes, monitoring — is operations, and it’s where the day-to-day risk lives.
People are the variable
The single largest cause of avoidable downtime is human error during change and routine work. Procedures (MOP/SOP/EOP), training and supervision are the controls that turn a good environment into a reliable one.
The risk is in the gaps
Outages cluster around the seams: an undocumented change, an un-diverse cross-connect, a monitoring blind spot, a single engineer who knows how the rack is wired. Operations is the discipline of closing those gaps before they find you.
Recovery is a capability
Failures are inevitable; extended outages are not. Whether a fault is a 30-second blip or a six-hour incident is decided by monitoring, escalation and rehearsed response — long before the alarm sounds.
The Operational Stack
The facility hands you power, cooling and space. Operations is everything done on top of that to keep your estate available — the hardware, the connectivity, the capacity, the monitoring and the people. The estate is only as available as the weakest of these, so each is a domain an operations team runs around the clock.
Hardware & rack operations
Installs, moves, swaps, break-fix and RMA handling. The hands-on work of bringing kit in, keeping live hardware healthy, and replacing what fails — correctly and to schedule.
Connectivity & cabling
Cross-connects, patching, port management, labelling and documentation. Power keeps kit alive; the network makes it useful — and undisciplined cabling is a slow-motion outage waiting to happen.
Capacity & asset management
Tracking space, power and port headroom against an accurate, live asset register — so you never strand capacity or, worse, run out of it mid-deployment.
Monitoring & alerting
Watching the environmental and power telemetry the facility exposes — temperature, humidity, load, access — and acting on it. Operations spots the deviation and escalates while there’s still time.
Physical security & access
Access control, escorting, CCTV review and audit logging. Who is on the floor and what they touch is part of availability and of compliance — and a discipline in its own right.
Single points of failure
The job is to find the link that isn’t resilient — the un-diverse cross-connect, the undocumented dependency, the single competent engineer. Unmapped SPOFs are where “resilient” estates still fail.
How Operationally Mature Are You?
Operational maturity is the gap between a site that has good kit and one that is run well. Score yourself honestly against the disciplines below — the bars show where most outages are won or lost. Anything you can’t evidence with a document or a record is a gap, not a strength.
Monitoring and procedures sit at the top deliberately — you can’t respond to what you can’t see, and you can’t act safely without a method. The most common failure pattern isn’t missing redundancy; it’s redundancy that was never tested, or a procedure that existed only in one engineer’s head.
What the “Nines” Really Cost
Availability is quoted in “nines,” and the difference between them is brutal: each extra nine cuts your allowable downtime roughly tenfold. Knowing what a target actually permits — in minutes per year — is the only honest way to decide what resilience you need and what it’s worth paying for.
| Availability | Downtime / year | Downtime / month | Typical context |
|---|---|---|---|
| 99% (“two nines”) | 3.65 days | 7.2 hrs | Non-critical / dev & test |
| 99.9% (“three nines”) | 8.76 hrs | 43.8 min | Standard business systems |
| 99.982% | 1.6 hrs | ~8 min | Common enterprise availability target |
| 99.99% (“four nines”) | 52.6 min | 4.4 min | Critical production |
| 99.999% (“five nines”) | 5.26 min | 26 sec | Mission-critical / fault tolerant |
These are the maths of availability, not a guarantee — a target only holds if the operations behind it do. “Five nines” means roughly five minutes of unplanned downtime in an entire year, which takes a resilient facility and, just as importantly, mature operations — and operations is the part most under-invested in. Note too the distinction between planned and unplanned downtime: a well-run estate schedules disruptive work into change-controlled maintenance windows, so the goal becomes zero unplanned outage.
MTBF — how often
Mean Time Between Failures: the expected interval between faults on a component. Higher is better, and it’s improved by quality kit and preventive maintenance.
MTTR — how fast
Mean Time To Repair/Recover: how long to restore service once something fails. This is the number operations controls most directly — through spares, monitoring and rehearsed response.
Planned ≠ unplanned
Disruptive work scheduled into a change-controlled maintenance window is planned and managed. Counting that as “downtime” conflates it with the real enemy — the goal is zero unplanned outage.
The Operational Disciplines
Reliable sites aren’t reliable by accident. They run a set of repeatable disciplines that catch problems early, make every action safe and repeatable, and keep the facility’s capacity ahead of demand. These are the practices that turn a redundant design into a resilient operation.
Change control
No change to a live system without review, approval, a method statement and a rollback plan. The single highest-leverage control against self-inflicted outages — because most outages are self-inflicted.
MOP / SOP / EOP
Method, Standard and Emergency Operating Procedures: the written, rehearsed scripts for routine work and for the bad day. When the alarm goes at 3am, you follow the EOP — you don’t improvise.
Capacity management
Tracking space, power and port headroom so you never strand capacity or, worse, run out of it mid-deployment. Stranded capacity is wasted spend; exhausted capacity is a stalled project.
Asset lifecycle management
Knowing what every asset is, where it sits and where it is in its life — from receipt and deployment through refresh to secure decommission. The backbone of capacity, support and audit.
Connectivity & cabling discipline
Structured cabling, change-managed patching, labelling and accurate records. Tidy, documented connectivity is fast to work on and safe to change; the alternative is an outage nobody can trace.
DCIM & documentation
An accurate, live record of every asset, circuit and connection. Decisions made on stale data cause outages; DCIM and an up-to-date asset register are the operational source of truth.
When something goes wrong: the incident lifecycle
Detection-to-recovery is a rehearsed sequence, not a scramble. Every minute saved here is a minute off your MTTR — and the difference between a near-miss and a reportable outage.
Detect
Monitoring catches the deviation — a temperature trend, a failed unit, a power anomaly — ideally before it affects the load.
Alert & triage
The alarm reaches the right person on a clear escalation path, and severity is assessed against defined thresholds.
Respond to procedure
The relevant EOP/SOP is executed — isolate, fail over, switch to alternate path — methodically, not improvised.
Stabilise
Service is protected or restored, the fault is contained, and the site is returned to a known-good, redundant state.
Restore & verify
The failed component is repaired or replaced, redundancy is re-established, and the fix is verified under load.
Review
A blameless post-incident review feeds back into procedures, training and design — so the same fault can’t bite twice.
How Optronix Helps
Operations is the longest stage of the data centre lifecycle — the years between building a facility and decommissioning it — and the one where reputation is made or lost. We support clients across the operational disciplines, with smart hands and live-operations cover, across the UK, Europe and worldwide — whether you run the site yourself or want us carrying the load.
Operational readiness & audit
We assess your estate against the disciplines above — procedures, change control, monitoring, capacity, cabling and SPOFs — and give you a prioritised plan to close the gaps.
Smart hands & remote ops
Eyes and hands on the floor for the lights-out site: installs, swaps, audits and break-fix, to your procedures and to schedule.
24/7 monitoring & response
Round-the-clock cover with clear escalation and rehearsed incident response — so an alarm at 3am is met with a procedure, not a scramble.
Capacity, asset & cabling support
Capacity tracking, asset-register and DCIM accuracy, and structured-cabling work — keeping your estate ahead of demand and audit-ready.
Running a facility — or stretched thin running it?
Tell us where it hurts — coverage gaps, an upcoming maintenance window, a site that needs eyes on it — and we’ll put the right operational support around it, on the ground wherever it is.
Sources & notes
- Uptime Institute Annual Outage Analysis — on human error as a factor in the majority of significant outages, and the planned/unplanned distinction
- Availability / downtime figures are the standard arithmetic of the “nines” (e.g. 99.999% ≈ 5.26 min/year)
- Common industry practice for change control, MOP/SOP/EOP, DCIM, and capacity & asset management
General guidance, not a substitute for site-specific engineering or your facility’s own procedures. Availability figures are illustrative of the maths, not a guarantee; achievable uptime depends on design, kit and — above all — operations. Confirm requirements against the relevant standards and your own risk appetite.