Operations Guide

Managed operations: tiers, response times & what an SLA really commits to

Most estates are not run by the people who own them. Somebody covers the floor, holds the runbook and answers the alarm at 3am — and how well that is defined decides whether an incident is an inconvenience or an outage. This guide covers what operations actually spans, how to read a response SLA properly, the five-tier model the market has largely settled on, and the arithmetic that decides which response time your uptime target can genuinely support.

Free, no email gate, with a maturity assessment you can score yourself against. The engineering and the arithmetic here are vendor-neutral and apply whoever you buy from; the tier model in section 06 is how Optronix delivers the service.
01 | Why It Matters

The Difference Is Almost Never the Kit

A hall is commissioned once and operated for a decade. Redundant power, N+1 cooling and a second carrier buy you the capacity to survive a fault — they do not decide whether you actually do. That is settled later, by whoever is on shift, what they are allowed to do, how quickly they can reach the rack, and whether the procedure they follow was ever tested. Two identical halls, built to the same drawing, post very different availability figures after three years, and the difference is almost never the kit.

Start here

This guide is about buying that operation rather than running it: what the scope really covers, how to read a response SLA, the tier model the market has settled on, and the arithmetic that decides which response time your uptime target can support. If you want the underlying engineering discipline instead — the availability stack, MTBF and MTTR, change control and the incident lifecycle in full — that is covered in running a data centre: operations, resilience & uptime.

02 | Scope

What “Operations” Actually Covers

Operations is often bought as a single line item and then discovered to be six. It is worth being explicit about scope before comparing anyone’s price, because the cheapest quote is almost always the one covering the least of it.

Smart handsSold by the fifteen minutes, on request
Hardware & rackConnectivity & cablingCapacity & assetsMonitoring & alertingSecurity & accessIncident ownership
Managed operationsHeld continuously, whether or not anything has broken
Hardware & rackConnectivity & cablingCapacity & assetsMonitoring & alertingSecurity & accessIncident ownership
covered    on request, chargeable    not covered
This is the distinction that decides the price, and it is why comparing the two on a rate card gets the wrong answer. Smart hands is one of these six, bought by the visit. Managed operations is all six, held between visits as well as during them — which is the part that stops the other five degrading quietly.
Hardware & rack

Installs, moves, adds, changes, goods receipt and RMA

Connectivity & cabling

Patching, cross-connects, labelling and record accuracy

Capacity & assets

Power, cooling, space and U tracked against what is installed

Monitoring & alerting

Telemetry with thresholds somebody owns

Security & access

Who may enter, with what authority and what audit trail

Incident ownership

Someone named who owns the fault through to closure

03 | Self-Assessment

How Operationally Mature Are You?

Operational maturity is the gap between a site that has good kit and one that is run well. Score each discipline honestly — anything you cannot evidence with a document or a record is a gap, not a strength. The disciplines are weighted by how often they decide the outcome of an incident, so monitoring and procedure count for more than training does.

Score each one — pick the option that you could actually evidence

Nothing is sent anywhere. The scoring runs in your browser and nothing is stored.

Monitoring and procedures are weighted highest deliberately — you cannot respond to what you cannot see, and you cannot act safely without a method. The most common failure pattern is not missing redundancy; it is redundancy that was never tested, or a procedure that existed only in one engineer’s head.

04 | Incidents

What Happens When It Goes Wrong

Every discipline above is really preparation for six minutes of someone else’s bad morning. This is the shape of those minutes, and the step a response SLA is actually measuring is the third one.

01Detect

Automatically and within seconds, not by a user reporting it. Detection you do not own caps everything after it.

02Alert & triage

Routed to someone awake, qualified and authorised, with severity assigned against a definition rather than a feeling.

03Respond

To procedure, with the access, tooling and spares already arranged. This is the step a response SLA measures.

04Stabilise

Stop it getting worse before making it better. Restoring service and fixing root cause are different jobs.

05Restore & verify

Service back and proven back, measured rather than assumed, with redundancy restored before closure.

06Review

Blameless, documented and turned into a change. An incident that alters no procedure will happen again.

05 | Service Levels

How to Read an Operations SLA

Response times are the most quoted and least examined number in this market. Four questions separate a commitment that will hold from one that reads well in a proposal.

Is it time to hands-on, or time to fixed?

Almost always the former, and it should say so. A response target is the clock until someone qualified is at the rack — diagnosis and repair come after it, out of the same outage budget. A response SLA quietly presented as a fix SLA is the single most common overclaim in operations contracts.

Where are the spares?

Fifteen minutes to a rack you cannot fix is fifteen minutes. A fast response is only worth what the on-site spares inventory makes it worth, so ask what is held, where, and who funds it. If the answer is a next-day courier, the real response time is next day.

How many people are actually on shift?

One technician on site means the second concurrent incident breaches by definition, and lone working is not permitted on live electrical or liquid-cooled plant anyway. Any genuine round-the-clock presence is a minimum of two people per shift — which is roughly eleven full-time staff once cover, leave and shift leads are counted.

What is committed at every tier, regardless?

Detection and notification should not vary by what you are paying. A P1 detected, notified and ticketed inside fifteen minutes is reasonable to expect at the lowest tier as well as the highest. The tier should change what happens next, not how quickly you find out.

The test

Ask a prospective provider what happens at 02:00 on a bank holiday when two racks fail at once. The answer exposes the shift pattern, the spares, the escalation path and the honesty of the response figure in a single question.

06 | The Model

The Optronix Tier Model: On-Call to Embedded

Optronix delivers managed operations on a five-step ladder. What separates the steps is not response time alone — it is presence (how often somebody is physically on your floor) and depth (how much of the stack we own). Response time follows from those two, which is why buying a fast response without the presence to support it does not work, from us or from anyone else.

01Copper

On-Call

No standing presence. Optronix attends on request, with your runbook and asset register already held and maintained by us between visits.

On siteOn request
Response48 hr / NBD
02Bronze

Scheduled

Fixed Optronix attendance days per site, working a managed ticket queue, with reactive cover in between.

On siteSet days per week
Response48 hr
03Silver

Managed

Two or more days a week on site plus round-the-clock monitoring, with Optronix owning the incident through to closure.

On site2+ days per week
Response8 hr
04Gold

Resident

Dedicated Optronix technicians on your floor every day across an extended two-shift pattern, on call overnight.

On siteDaily, two shifts
Response30 min
05Platinum

Embedded

A dedicated Optronix crew on site around the clock, with spares held and a committed availability target. Optronix owns your DC operations function outright — end to end, not shared with your team.

On site24/7/365
Response15 min
The span is deliberately wide, and both ends of it are real services. At one end this is light touch: nobody standing on your floor, a runbook and an asset register kept current, and an engineer dispatched when you actually need one. At the other it is a fully outsourced operations team — technicians on site 24/7/365, owning the estate end to end, with spares held on site and a committed availability target. Same model, same portal, same reporting either way. The only variable is how much of the operation you want to keep holding yourself.

Tiers 01 to 03 are priced on volume; 04 and 05 are priced on a rota, because they are staffing commitments rather than task commitments. That is also why the top two normally carry a minimum term and a mobilisation fee — nobody hires eleven people against a rolling monthly contract.

What tends to sit at each step

Capability01On-Call02Scheduled03Managed04Resident05Embedded
Installs, moves, adds and changes●●●●●
Ticketing portal, managed queue and service reporting●●●●●
Goods receipt, shipping and RMA logistics○●●●●
Asset register ownership and audit○●●●●
Photographic evidence on every change○●●●●
Round-the-clock infrastructure monitoring○○●●●
Incident ownership through to closure○○●●●
OEM and warranty escalation management○○●●●
Named service delivery manager○○●●●
Live DCIM for power, thermal and capacity○○○●●
On-site spares pool held and managed○○○○●
Committed availability target○○○○●

● included as standard    ○ available as an option. The ladder is a starting point, not a limit — any capability can be added to any tier, and the mix can differ site by site.

07 | The Systems

What You Actually Log Into

Every tier includes the portal: raise and track tickets, interrogate the asset register and pull service reports on demand. From the upper tiers we instrument the estate itself, so power, thermal and capacity arrive live in DCIM rather than being asked for. These are our own systems, built in house rather than licensed in — which is the reason they can be shaped around the way you already work.

We are not precious about tooling. Run ours standalone; keep what you already have and we operate it; or run both, with two-way ticket sync over API so an event on the floor opens, updates and closes in your system without anyone typing. Taking on an incumbent’s tooling is a normal part of a transition, not an exception to it.
portal.optronix.co.uk/dashboard Dashboard

Uptime, power, PUE, thermals and open work across the estate, on one screen.

What you’d use it for
portal.optronix.co.uk/tickets Service desk

Every ticket, its SLA clock, and the full activity trail from the floor.

What you’d use it for
portal.optronix.co.uk/dcim DCIM

Live rack elevations, occupancy, draw and inlet temperature, rack by rack.

What you’d use it for

Live previews of the real systems, running in the page. Select any one of them to see what it is used for.

08 | The Allowance

How Much Downtime Your Target Allows

Availability targets are quoted far more often than they are costed. What the figure actually buys is a fixed allowance of unplanned downtime a year, and every minute of it is spent whether somebody is travelling to site or already working on the fault. It is worth seeing the numbers written down before agreeing to one.

Unplanned downtime per year, to scale
99%3 days 15 hrA single bad day spends the whole year.
99.5%1 day 20 hrOne long incident takes most of the year with it.
99.9%8 hr 46 minA next-day attendance no longer fits inside one outage.
99.95%4 hr 23 minSame-day attendance is the floor, and travel time is now material.
99.99%52 minUnder an hour a year. Somebody has to already be on the floor.
99.999%5 min 15 secWon with redundancy and concurrent maintainability, not with response.

Planned maintenance is excluded — most availability targets are written against unplanned outage only, and it is worth checking which yours means before committing to it.

Read the last column against the response times in the tier table above, and remember that a response clock is spent before any diagnosis starts. Around three nines a same-day attendance still fits inside a single outage. By four it does not: the travel alone would spend the year’s budget, so the only way to buy the response is to have somebody already standing on the floor. Beyond that, availability stops being an operations question at all — five nines is 5 minutes a year, which is won with redundancy and concurrent maintainability. Operations is what keeps it there.
09 | Getting Started

The Two Ways an Operations Contract Begins

Whoever you appoint, there are only two starting positions: an estate that does not exist yet, and one that is already running. They are different pieces of work, and the second is the one most often underestimated.

From nothing — deploy and commission

Site selection, goods receipt, structured cabling, racking (or a migration of existing kit), labelling, staged power-on, burn-in, validation and a formal handover into steady state. Normally priced as a project rather than a subscription. The advantage is that whoever runs it afterwards already knows every cable in it.

From a live floor — survey, integrate, take over

A physical survey against the existing records, a baselined asset register, monitoring and DCIM connected, then the runbook, escalation path and SLA formally assumed. Nothing is rebuilt and nothing stops. The work is in the record: almost every live estate has drifted from its documentation, and the audit is what makes the SLA safe to sign.

Watch for

If you are moving between providers, the transition audit is the part to scrutinise. An incoming operator who accepts an inherited asset register without verifying it is accepting your drift as their baseline — and the first incident is where that gets discovered.

10 | The Bridge

How Optronix Helps

Operations is the longest stage of the data centre lifecycle — the years between commissioning a facility and decommissioning it — and the one where reputation is made or lost. We run the model set out above, at either end of it: called to site only when we are needed, or a full Optronix crew on your floor around the clock owning the whole operation. Five tiers on presence and depth, one portal, one ticket queue and a named service manager, whether we built the floor or inherited it.

Your presence on site

Not remote hands sold by the fifteen minutes. We become the hands on the rack, the eyes on the estate and the system of record for every asset in it — so your team never needs to enter the facility.

Any tier, any number of sites

From attendance on request through to a dedicated crew on the floor around the clock. Scheduled operations are running today across five cities in two continents, and every tier is available anywhere we can reach and staff a site.

The floor and the fabric

Rack-level operations and network operations run on the same ladder and the same queue, so an alert on a switch and a body in the hall are one ticket rather than two suppliers.

Evidence, not assurances

Photographic evidence on every change, a maintained asset register, live DCIM from the upper tiers, and reporting you can hand to your own auditors.

If you would rather see the whole model first, the deck runs through every tier, the capability matrix and the portal — and the systems above are running live inside it.

Managed operations.Service pack

Managed Operations

The five-tier service ladder for the floor and the fabric — take one, the other or both. Capability matrix per tier, what a fifteen-minute response really commits to, and the portal running live in the deck.

23 slides15 min
Open deck

Which tier does your estate actually need?

Tell us the site, the coverage you need and the hours that matter, and we will come back with the tier that fits.

Talk to our team

Sources & notes

  • Uptime Institute Annual Outage Analysis — on human error as a factor in the majority of significant outages, and the planned/unplanned distinction
  • Availability and downtime figures are the standard arithmetic of the “nines” (e.g. 99.999% ≈ 5.26 min/year)
  • Common industry practice for change control, MOP/SOP/EOP, DCIM, and capacity & asset management
  • The five-tier structure in section 06 is the Optronix service model. Tier names are descriptive; there is no industry standard for them