Managed operations: tiers, response times & what an SLA really commits to
Most estates are not run by the people who own them. Somebody covers the floor, holds the runbook and answers the alarm at 3am — and how well that is defined decides whether an incident is an inconvenience or an outage. This guide covers what operations actually spans, how to read a response SLA properly, the five-tier model the market has largely settled on, and the arithmetic that decides which response time your uptime target can genuinely support.
The Difference Is Almost Never the Kit
A hall is commissioned once and operated for a decade. Redundant power, N+1 cooling and a second carrier buy you the capacity to survive a fault — they do not decide whether you actually do. That is settled later, by whoever is on shift, what they are allowed to do, how quickly they can reach the rack, and whether the procedure they follow was ever tested. Two identical halls, built to the same drawing, post very different availability figures after three years, and the difference is almost never the kit.
This guide is about buying that operation rather than running it: what the scope really covers, how to read a response SLA, the tier model the market has settled on, and the arithmetic that decides which response time your uptime target can support. If you want the underlying engineering discipline instead — the availability stack, MTBF and MTTR, change control and the incident lifecycle in full — that is covered in running a data centre: operations, resilience & uptime.
What “Operations” Actually Covers
Operations is often bought as a single line item and then discovered to be six. It is worth being explicit about scope before comparing anyone’s price, because the cheapest quote is almost always the one covering the least of it.
Installs, moves, adds, changes, goods receipt and RMA
Patching, cross-connects, labelling and record accuracy
Power, cooling, space and U tracked against what is installed
Telemetry with thresholds somebody owns
Who may enter, with what authority and what audit trail
Someone named who owns the fault through to closure
How Operationally Mature Are You?
Operational maturity is the gap between a site that has good kit and one that is run well. Score each discipline honestly — anything you cannot evidence with a document or a record is a gap, not a strength. The disciplines are weighted by how often they decide the outcome of an incident, so monitoring and procedure count for more than training does.
Nothing is sent anywhere. The scoring runs in your browser and nothing is stored.
Monitoring and procedures are weighted highest deliberately — you cannot respond to what you cannot see, and you cannot act safely without a method. The most common failure pattern is not missing redundancy; it is redundancy that was never tested, or a procedure that existed only in one engineer’s head.
What Happens When It Goes Wrong
Every discipline above is really preparation for six minutes of someone else’s bad morning. This is the shape of those minutes, and the step a response SLA is actually measuring is the third one.
Automatically and within seconds, not by a user reporting it. Detection you do not own caps everything after it.
Routed to someone awake, qualified and authorised, with severity assigned against a definition rather than a feeling.
To procedure, with the access, tooling and spares already arranged. This is the step a response SLA measures.
Stop it getting worse before making it better. Restoring service and fixing root cause are different jobs.
Service back and proven back, measured rather than assumed, with redundancy restored before closure.
Blameless, documented and turned into a change. An incident that alters no procedure will happen again.
How to Read an Operations SLA
Response times are the most quoted and least examined number in this market. Four questions separate a commitment that will hold from one that reads well in a proposal.
Is it time to hands-on, or time to fixed?
Almost always the former, and it should say so. A response target is the clock until someone qualified is at the rack — diagnosis and repair come after it, out of the same outage budget. A response SLA quietly presented as a fix SLA is the single most common overclaim in operations contracts.
Where are the spares?
Fifteen minutes to a rack you cannot fix is fifteen minutes. A fast response is only worth what the on-site spares inventory makes it worth, so ask what is held, where, and who funds it. If the answer is a next-day courier, the real response time is next day.
How many people are actually on shift?
One technician on site means the second concurrent incident breaches by definition, and lone working is not permitted on live electrical or liquid-cooled plant anyway. Any genuine round-the-clock presence is a minimum of two people per shift — which is roughly eleven full-time staff once cover, leave and shift leads are counted.
What is committed at every tier, regardless?
Detection and notification should not vary by what you are paying. A P1 detected, notified and ticketed inside fifteen minutes is reasonable to expect at the lowest tier as well as the highest. The tier should change what happens next, not how quickly you find out.
Ask a prospective provider what happens at 02:00 on a bank holiday when two racks fail at once. The answer exposes the shift pattern, the spares, the escalation path and the honesty of the response figure in a single question.
The Optronix Tier Model: On-Call to Embedded
Optronix delivers managed operations on a five-step ladder. What separates the steps is not response time alone — it is presence (how often somebody is physically on your floor) and depth (how much of the stack we own). Response time follows from those two, which is why buying a fast response without the presence to support it does not work, from us or from anyone else.
On-Call
No standing presence. Optronix attends on request, with your runbook and asset register already held and maintained by us between visits.
Scheduled
Fixed Optronix attendance days per site, working a managed ticket queue, with reactive cover in between.
Managed
Two or more days a week on site plus round-the-clock monitoring, with Optronix owning the incident through to closure.
Resident
Dedicated Optronix technicians on your floor every day across an extended two-shift pattern, on call overnight.
Embedded
A dedicated Optronix crew on site around the clock, with spares held and a committed availability target. Optronix owns your DC operations function outright — end to end, not shared with your team.
Tiers 01 to 03 are priced on volume; 04 and 05 are priced on a rota, because they are staffing commitments rather than task commitments. That is also why the top two normally carry a minimum term and a mobilisation fee — nobody hires eleven people against a rolling monthly contract.
What tends to sit at each step
| Capability | 01On-Call | 02Scheduled | 03Managed | 04Resident | 05Embedded |
|---|---|---|---|---|---|
| Installs, moves, adds and changes | ● | ● | ● | ● | ● |
| Ticketing portal, managed queue and service reporting | ● | ● | ● | ● | ● |
| Goods receipt, shipping and RMA logistics | ○ | ● | ● | ● | ● |
| Asset register ownership and audit | ○ | ● | ● | ● | ● |
| Photographic evidence on every change | ○ | ● | ● | ● | ● |
| Round-the-clock infrastructure monitoring | ○ | ○ | ● | ● | ● |
| Incident ownership through to closure | ○ | ○ | ● | ● | ● |
| OEM and warranty escalation management | ○ | ○ | ● | ● | ● |
| Named service delivery manager | ○ | ○ | ● | ● | ● |
| Live DCIM for power, thermal and capacity | ○ | ○ | ○ | ● | ● |
| On-site spares pool held and managed | ○ | ○ | ○ | ○ | ● |
| Committed availability target | ○ | ○ | ○ | ○ | ● |
● included as standard ○ available as an option. The ladder is a starting point, not a limit — any capability can be added to any tier, and the mix can differ site by site.
What You Actually Log Into
Every tier includes the portal: raise and track tickets, interrogate the asset register and pull service reports on demand. From the upper tiers we instrument the estate itself, so power, thermal and capacity arrive live in DCIM rather than being asked for. These are our own systems, built in house rather than licensed in — which is the reason they can be shaped around the way you already work.
Uptime, power, PUE, thermals and open work across the estate, on one screen.
What you’d use it forEvery ticket, its SLA clock, and the full activity trail from the floor.
What you’d use it forLive rack elevations, occupancy, draw and inlet temperature, rack by rack.
What you’d use it forLive previews of the real systems, running in the page. Select any one of them to see what it is used for.
How Much Downtime Your Target Allows
Availability targets are quoted far more often than they are costed. What the figure actually buys is a fixed allowance of unplanned downtime a year, and every minute of it is spent whether somebody is travelling to site or already working on the fault. It is worth seeing the numbers written down before agreeing to one.
Planned maintenance is excluded — most availability targets are written against unplanned outage only, and it is worth checking which yours means before committing to it.
The Two Ways an Operations Contract Begins
Whoever you appoint, there are only two starting positions: an estate that does not exist yet, and one that is already running. They are different pieces of work, and the second is the one most often underestimated.
From nothing — deploy and commission
Site selection, goods receipt, structured cabling, racking (or a migration of existing kit), labelling, staged power-on, burn-in, validation and a formal handover into steady state. Normally priced as a project rather than a subscription. The advantage is that whoever runs it afterwards already knows every cable in it.
From a live floor — survey, integrate, take over
A physical survey against the existing records, a baselined asset register, monitoring and DCIM connected, then the runbook, escalation path and SLA formally assumed. Nothing is rebuilt and nothing stops. The work is in the record: almost every live estate has drifted from its documentation, and the audit is what makes the SLA safe to sign.
If you are moving between providers, the transition audit is the part to scrutinise. An incoming operator who accepts an inherited asset register without verifying it is accepting your drift as their baseline — and the first incident is where that gets discovered.
How Optronix Helps
Operations is the longest stage of the data centre lifecycle — the years between commissioning a facility and decommissioning it — and the one where reputation is made or lost. We run the model set out above, at either end of it: called to site only when we are needed, or a full Optronix crew on your floor around the clock owning the whole operation. Five tiers on presence and depth, one portal, one ticket queue and a named service manager, whether we built the floor or inherited it.
Your presence on site
Not remote hands sold by the fifteen minutes. We become the hands on the rack, the eyes on the estate and the system of record for every asset in it — so your team never needs to enter the facility.
Any tier, any number of sites
From attendance on request through to a dedicated crew on the floor around the clock. Scheduled operations are running today across five cities in two continents, and every tier is available anywhere we can reach and staff a site.
The floor and the fabric
Rack-level operations and network operations run on the same ladder and the same queue, so an alert on a switch and a body in the hall are one ticket rather than two suppliers.
Evidence, not assurances
Photographic evidence on every change, a maintained asset register, live DCIM from the upper tiers, and reporting you can hand to your own auditors.
If you would rather see the whole model first, the deck runs through every tier, the capability matrix and the portal — and the systems above are running live inside it.
Which tier does your estate actually need?
Tell us the site, the coverage you need and the hours that matter, and we will come back with the tier that fits.
Sources & notes
- Uptime Institute Annual Outage Analysis — on human error as a factor in the majority of significant outages, and the planned/unplanned distinction
- Availability and downtime figures are the standard arithmetic of the “nines” (e.g. 99.999% ≈ 5.26 min/year)
- Common industry practice for change control, MOP/SOP/EOP, DCIM, and capacity & asset management
- The five-tier structure in section 06 is the Optronix service model. Tier names are descriptive; there is no industry standard for them
Sections 01 to 05 are general guidance and apply whoever you buy from. They are not a substitute for site-specific engineering or your facility’s own procedures. Downtime figures are the arithmetic of the nines, not a guarantee: real availability depends on design, kit and — above all — operations. Confirm requirements against the relevant standards, your contracts and your own risk appetite.