01 / The challengeWhy hyperscale deployment is different
A hyperscale hall is built to a capacity date, and almost every decision about how the deployment runs follows from that.
The date is committed long before the hall is ready for the installers, the link count runs into the tens of thousands, and the same work is repeated across hundreds of cabinets by different crews on different shifts, weeks apart. Four things make it a different job from an enterprise fit-out.
The date is fixed
Capacity is committed before the first pod is released. The last wave carries no float, so the programme is built backwards from the day the pods have to be live, with waves sized to the access the site can actually give.
The link count
One pod carries thousands of terminations across the compute, storage and management fabrics. A single contaminated endface in that count stays quiet until the fabric is under load.
There is no second pass
Once nodes are in service the row is out of reach. Work that should have been finished on the day the cabinet was racked becomes a change request with an outage attached, so quality is checked inside the crew, on the day.
Identical, everywhere
Different crews, different shifts, weeks apart, and the same build every time. Consistency at that scale comes from a written method and in-crew checks rather than from individual judgement.
02 / The hallNine layers, one section through the row
A hyperscale hall is easiest to understand as a section through one row: where the fibre arrives, how it is distributed, the spine and leaf layers, the compute pods that repeat across the hall, and the storage, management, pathway and power layers that serve them. Each layer is built, and recorded, in its own way.
Layer 05
GPU compute pods
The repeating scale unit. Rows, pods, cabinets per pod and nodes per cabinet, built the same way every time.
The compute pod is the scale unit. Once its design is right, the hall is that pod repeated, so most of a deployment method is written around building one pod the same way every time, and around the shared layers (distribution, spine, pathway and power) that have to be in and proven before the first pod is loaded.
03 / The fabricA rail-optimised pod
Inside a GPU pod there are several networks, not one, and the cabling for each is planned separately.
The compute fabric carries the east-west traffic between accelerators and is usually wired rail-optimised: each accelerator in a node connects to a different leaf switch, so the eight accelerators of a node land on eight different rails. A leaf failure then costs one rail across the pod rather than every port of a node. Storage runs on its own fabric with its own leaves, out-of-band management has its own switches and pathway, and every leaf has uplinks to every spine.
The rails
Compute east-west
One port per accelerator, and every port on a different leaf. Eight rails across the pod, so a leaf failure costs one rail and never a whole node.
Eight of thirty-two nodes shown. Hover or tap a leaf to follow its rail. Pod size, rail count and port speeds come from the reference design being built.
The rail design is why the port maps matter so much on site. A lead in the wrong leaf port still lights, still passes a basic continuity check and quietly breaks the rail. Every port is mapped to a node and a rail before the first lead is made up, and the map is checked again at test.
04 / Speed steps100G to 1.6T on one physical layer
800G is shipping in hyperscale halls today and 1.6T is next. The lane count is what drives the connector, so the trunking that goes in now is specified for the step after the one being bought.
Parallel singlemode optics carry a transmit and a receive fibre per lane. At 100G and 400G that is four lanes and eight fibres, which sit in a 12-fibre MPO with four positions unused. At 800G the lane count doubles to eight, and sixteen fibres need a 16-fibre connector. 1.6T doubles the lane rate rather than the lane count, so it runs on the same sixteen fibres.
The lane count doubles
800G
Eight lanes, sixteen fibres. This is the step that changes the connector, and the reason to put base-16 trunking in before it is needed.
- The optics
- The patch leads
- The trunk, if base-12 went in
- Base-16 trunks and cassettes
- Containment and pathway
- Labelling and as-builts
Parallel singlemode shown, because that is what carries the lane count up the curve. Duplex wavelength-multiplexed variants exist at every step for longer reach. Reaches are the usual figures for each optic class, not a guarantee for a particular part.
A hall cabled in base-12 for 400G has to re-pull its trunks for 800G, or run conversion modules that strand fibres and add mated pairs. Base-16 trunks and cassettes installed from the start make 800G and 1.6T a change of optics and patch leads. MPO-16 is keyed differently from MPO-12, so the two cannot be mated by accident.
05 / The link budgetEvery mated pair spends budget
At 800G the optic’s loss allowance is small, and it does not grow with the speed. Every connector in the channel spends part of it.
A link from a node in one pod to a spine in the main distribution can pass through six mated pairs: patch cord to cassette, through the trunk, through the distribution frames and out again. With standard-grade connectivity at 0.5 dB a pair, six pairs and 300 m of singlemode already exceed the 3.0 dB allowance of an eight-lane parallel optic. The same channel built from ultra-low-loss components passes with margin.
Illustrative planning figures. Every link is designed against the published channel specification of the optic being fitted.
This is why connector grade and cleanliness are design decisions rather than installation details. Ultra-low-loss components buy the margin that a contaminated endface, an extra patch or a re-route would otherwise use up, and at 1.6T the allowance is no larger than it was at 800G.
06 / Pre-terminatedPre-terminated, because the density demands it
At hyperscale link counts, field termination is the slowest and least repeatable part of the job.
Trunks and cassettes arrive terminated and measured, polarity is a scheme rather than a decision taken at each link, and the hall can be re-lit at the next speed step by changing patch leads.
Terminated and tested off site
Trunks and cassettes are terminated and measured under factory conditions and arrive with a result per fibre. Consistent loss across thousands of links comes from a controlled process.
Polarity as a scheme
One polarity method is chosen for the hall and carried through every trunk, cassette and lead. Nobody works it out at the cabinet, and a lead taken from stores is the right lead.
Inspect and clean, every mate
Every endface is inspected and cleaned to IEC 61300-3-35 before it is mated, including the ones that came out of a sealed bag. Contamination is the largest single cause of avoidable link failures at these densities.
Re-light, rather than re-pull
When the hall steps up a speed, the optics and the patch leads change. Base-16 trunks, cassettes, pathway and labelling stay exactly where they are, and the hall stays in service while it happens.
07 / The deployment lineGoods-in to handover, in one order
The same ten steps in the same order on every cabinet in the hall.
If a crew has to work out what comes next, cabinet six hundred will not match cabinet one. A written line, with the evidence each step produces, is what keeps a hall consistent across crews and shifts, and it is what lets a supervisor see at a glance where every cabinet stands.
Step 10 of 10
Handover
As-builts, patching schedules, the asset register and the full test results handed back in the operator’s own format, walked through at the hold point and signed off.
Hover or tap a step to hold it.
08 / The cabinetThe decisions inside one cabinet
A loaded GPU cabinet fills on power and weight long before it fills on rack units.
The elevation below is an example, drawn to show how the decisions inside one cabinet fit together: where the fibre field sits, how the switching and out-of-band are placed, how heavy nodes are handled and dressed, and why free rack units are not spare capacity. On a real deployment the elevation is the operator’s, and the cabinet is built to it U for U.
Detail 04
Compute nodes
Heavy, deep and awkward. The handling plan matters as much as the elevation.
- example frame
- 52U
- used in this elevation
- 40U
- blanked, not spare capacity
- 12U
- feeds to two rack PDUs
- A+B
An example elevation, drawn to show how the decisions fit together. Not to scale.
09 / The standardThe operator issues the standard
At this scale the operator owns the method. The installer’s job is to work inside it exactly, in every cabinet.
Six documents usually come from the operator, and each one has a matching discipline on the floor.
Issued by the operator
The rack elevation
Built U for U and checked against the drawing before the cabinet is signed off. Where the kit does not match the drawing, it is raised rather than absorbed.
Issued by the operator
The labelling scheme
Applied to every port, panel, lead and cabinet at both ends, and carried into the as-builts and the patching schedule without a local variation.
Issued by the operator
The routing and dressing spec
Followed to the letter, including the parts that only matter in year three: bundle sizes, tie spacing, radius control and which side a cord leaves on.
Issued by the operator
The inspection and test plan
Worked to as written, with the hold points held. Nothing moves past one because the programme is tight, and the witness is invited rather than assumed.
Issued by the operator
The evidence requirement
Photographs at the operator’s checkpoints, against the cabinet reference, uploaded from the floor rather than assembled afterwards.
Issued by the operator
The asset register format
Populated in the operator’s schema, with their field names and validation, so the file imports on the day it lands instead of coming back for a re-key.
The installer should propose one in writing, have it approved before the first cabinet, and hold it as the standard for the rest of the hall. Writing it down once is cheaper than discovering, at cabinet three hundred, that two crews have been dressing the same leads two different ways.
10 / Waves and crewsThe hall fills in waves
A hall is released in waves of pods, and the crew ramps up and down with them.
Each wave runs the same sequence: cabinet build, node install, fabric cabling, then test and evidence. Waves overlap, so in a typical week one wave is being cabled while the next is being built and the one before is under test. The crew size follows the overlap, and the programme is worked back from the date the first pods have to be live.
Week9of 18
2 pods certified and handed over
Wave 2, fabric cabling. Wave 3, cabinet build.
Wave 3, node install, week 10
- pods handed over
- 2
- engineers on site this week
- 34
An illustrative programme, drawn to show the shape of a two-pod wave, not a quote. Real wave sizes come from the release dates and the access the site can give.
11 / ScaleSizing the hall
The numbers that drive a deployment come from a small model: cabinets, nodes per cabinet, and the fabrics each node connects to.
Move the sliders to size a hall. Every figure below is built from the model shown, so you can see what drives it.
Each block is a pod of eight cabinets128 cabinets, 16 pods
- compute nodes
- 512
- ports on the compute fabric
- 4,096
- fibre links, every fabric
- 6,400
- copper links, out-of-band
- 896
- engineer-days to build, cable and dress
- 259
- days to certify 100% of the links
- 9
Model8 cabinets per pod8 accelerators per node2 storage links per node10 leaves per pod, 8 uplinks each1 out-of-band link per node, 3 per cabinet160 links per engineer-day3 nodes and 3 cabinets per engineer-day900 links per test-day
Illustrative planning figures. Every deployment is surveyed and scoped before a programme or a price is committed.
Two things stand out at any size. Fibre links outnumber nodes many times over, because every accelerator has its own port on the compute fabric; and certification, at around 900 links a test-day in this model, is a real line in the programme rather than a final afternoon.
12 / Test and evidence100% certified, link by link
At handover the evidence is the deliverable: a result for every link, a photograph set for every cabinet, and records that import on the day they land.
Fibre is tested for insertion loss and with an OTDR, with launch and receive cables fitted so the first and last connectors are measured. Copper out-of-band links are certified as permanent links to TIA-568 and ISO/IEC 11801. Nothing is sampled.
MDA-A03-U12-P01 to DH1-B07-U40-P01
DH1-B07-U45-P12 to DH1-B07-U18-P12
Facsimiles, built to show the shape of the record rather than a real result.
Every link, no sampling
- Fibre: insertion loss and OTDR, with launch and receive cables fitted on every link, and every endface inspected and cleaned before it is mated.
- Copper: channel and permanent link testing to TIA-568 and ISO/IEC 11801, on calibrated field testers.
- Defects: reported, remediated and retested before the link is signed off.
What exists at handover
- Test results, link by link, fibre and copper
- Photographs per cabinet, at the operator’s checkpoints
- As-built drawings and containment routes
- The patching schedule, both ends of every link
- The asset register, in the operator’s format
- Every deviation raised, with the approval against it
DH1-B07-U42-P24toMDA-A03-U10-P24
- DH1
- Data hall
- B07
- Row and cabinet
- U42
- Panel position
- P24
- Port
- Both ends
- Same rule at the far end
13 / How Optronix helpsOne owner for the fabric and the kit
Structured cabling, containment, power distribution and white-space fit-out are our core business. We design, supply, install, certify and migrate across the full data centre lifecycle for hyperscale, colocation and enterprise clients, across the UK, EMEA and Asia-Pacific.
Our engineers are ours
Optronix employs its engineers directly. The people on the hall are on our payroll and under our supervision, working to one written method on every shift.
Current at this density
A large share of our current work is AI and GPU deployment, where fibre density has moved to 800G. This is the work we are doing now, not work we are preparing for.
Evidence as standard
Certified links, photographs, as-builts and a populated asset register are part of the work, in the operator’s format, and not a line item we come back for.
One scope, from goods-in to handover
- 01
Receipt and asset capture
Deliveries checked against the manifest, damage recorded, and every serial, tag and position captured before a box leaves goods-in.
- 02
Cabinet and rail build
Cabinets set out to the floor grid, bayed, levelled and bonded, then rails and cable management built to the issued elevation.
- 03
Compute and GPU nodes
Heavy accelerated nodes handled with lifting aids and two people, seated on their rails, cabled clear of the airflow path and powered on in order.
- 04
Power to the rack
A and B feeds landed into the rack PDUs, outlets mapped to the elevation, cords cut to length, retained and dressed to the correct side.
- 05
Pathways and containment
Tray, basket and in-row managers installed and loaded to the design fill, with bend radius kept on every fibre route.
- 06
Fabric cabling
Compute, storage, in-band and out-of-band networks, installed as pre-terminated structured links with polarity managed as one scheme.
- 07
Labelling and records
Every port, panel and lead labelled at both ends to the operator’s scheme, and the asset register and patching schedule built as the work goes in.
- 08
Test, certify, evidence
100% of links certified, results recorded link by link, photographs per cabinet, and the pack handed back in the operator’s format.
- 01
Survey and plan
The reference design, the pod count and the first live date are enough to start: the scope, the crew shape and an outline programme come back worked from that date.
- 02
Mobilise and build
Goods-in, asset capture, cabinet build and node install, by directly employed engineers with a supervisor on every shift.
- 03
Cable, test, certify
Pre-terminated fabric cabling to the port maps, every endface inspected and cleaned, and 100% of links certified.
- 04
Evidence and handover
Records in the operator’s format, issued wave by wave and walked through at the hold points.
Related services
The physical layer is covered in depth in structured cabling; the power, cooling and fabric design of the halls these deployments fill in planning an AI/GPU-ready data hall; and the shell-to-live programme around them in the data centre fit-out.
Hyperscale deployment
Tell us the hall and the date.
The reference design, the pod count and the date the first pods have to be live are enough to start. We will come back with the scope, the crew shape and an outline programme worked back from that date.
Sources and notes
- TIA-942 and ISO/IEC 11801: data centre cabling spaces, the distribution hierarchy and media.
- IEEE 802.3: 400GBASE-DR4 and 800GBASE-DR8 parallel singlemode optics (802.3bs, 802.3df) and their channel insertion loss allowances; 1.6T on 200G lanes (802.3dj).
- IEC 61300-3-35: visual inspection of fibre optic connector endfaces.
- TIA-568.3 and the IEC 61280-4 series: optical fibre cabling test methods, insertion loss and OTDR, with launch and receive cables.
- TIA-568.2 and ISO/IEC 11801: balanced copper permanent link and channel testing.
- IEC 61754-7 family: MPO connector interfaces, including the 12-fibre and 16-fibre variants and their keying.
- General industry practice for rail-optimised GPU fabrics, pre-terminated base-16 cabling and wave-based hall deployment.
General guidance, not a design specification. Loss figures, reaches, lane rates, the wave programme and the planning model are illustrative and depend on the specific optics, components and site. Confirm against the published specifications and an engineered design.