Optronix, a part of the GS Group

16Technical guide

Hyperscale deployment: cabinet six hundred, built like cabinet one

A hyperscale hall is built to a capacity date. The date is committed before the first pod is released, the link count runs into the tens of thousands, and once a row is loaded there is no window to go back and do it again. This guide covers how a deployment at that scale runs: the hall and its fabrics, the physical layer from 400G to 1.6T, the link budget, the deployment line, the cabinet, the waves, and the evidence that has to exist at handover.

Vendor-neutral and educational. Standards and figures from TIA-942, ISO/IEC 11801, IEEE 802.3, IEC 61300-3-35 and published optic specifications; see the notes at the end.

A GPU compute pod in a hyperscale hallAn isometric drawing of one pod of eight cabinets with a second row behind, the overhead tray carrying the fibre tier, a drop into every cabinet, a leaf cabinet in the row, and a trunk leaving for the spine row. Pod 01 / 8 cabinets Fibre tier Base-16 trunks To spine row Rail leaves 800G / MPO-16 / APC Section / not to scale
Pod 01 / section
Dwg
OX-RES-16
Rev
4.0
Scale
NTS
shipping in hyperscale halls today, on eight lanes and sixteen fibres
800G
the next step, on the same base-16 physical layer (IEEE 802.3dj)
1.6T
per ultra-low-loss mated pair, half the standard-grade figure
0.25dB
of links certified and recorded, fibre and copper, with no sampling
100%

01 / The challengeWhy hyperscale deployment is different

A hyperscale hall is built to a capacity date, and almost every decision about how the deployment runs follows from that.

The date is committed long before the hall is ready for the installers, the link count runs into the tens of thousands, and the same work is repeated across hundreds of cabinets by different crews on different shifts, weeks apart. Four things make it a different job from an enterprise fit-out.

01

The date is fixed

Capacity is committed before the first pod is released. The last wave carries no float, so the programme is built backwards from the day the pods have to be live, with waves sized to the access the site can actually give.

02

The link count

One pod carries thousands of terminations across the compute, storage and management fabrics. A single contaminated endface in that count stays quiet until the fabric is under load.

03

There is no second pass

Once nodes are in service the row is out of reach. Work that should have been finished on the day the cabinet was racked becomes a change request with an outage attached, so quality is checked inside the crew, on the day.

04

Identical, everywhere

Different crews, different shifts, weeks apart, and the same build every time. Consistency at that scale comes from a written method and in-crew checks rather than from individual judgement.

02 / The hallNine layers, one section through the row

A hyperscale hall is easiest to understand as a section through one row: where the fibre arrives, how it is distributed, the spine and leaf layers, the compute pods that repeat across the hall, and the storage, management, pathway and power layers that serve them. Each layer is built, and recorded, in its own way.

DiagramA section through a hyperscale hallSelect a layer
Entrance roomDistributionManagementSpine rowLeaf and ToRGPU computeStorageShared distributionScale unit 01Scale unit 02 Fibre tray Out-of-band A and B feeds

Layer 05

GPU compute pods

The repeating scale unit. Rows, pods, cabinets per pod and nodes per cabinet, built the same way every time.

The compute pod is the scale unit. Once its design is right, the hall is that pod repeated, so most of a deployment method is written around building one pod the same way every time, and around the shared layers (distribution, spine, pathway and power) that have to be in and proven before the first pod is loaded.

03 / The fabricA rail-optimised pod

Inside a GPU pod there are several networks, not one, and the cabling for each is planned separately.

The compute fabric carries the east-west traffic between accelerators and is usually wired rail-optimised: each accelerator in a node connects to a different leaf switch, so the eight accelerators of a node land on eight different rails. A leaf failure then costs one rail across the pod rather than every port of a node. Storage runs on its own fabric with its own leaves, out-of-band management has its own switches and pathway, and every leaf has uplinks to every spine.

DiagramRail-optimised GPU podSelect a fabric
SpineSpineSpineSpineSpine row Rail leavesLeaf 1Leaf 2Leaf 3Leaf 4Leaf 5Leaf 6Leaf 7Leaf 8 Storage leaf AStorage leaf BOut-of-band switch N01N02N03N04N05N06N07N08 Eight of thirty-two nodes shown. Click a leaf to follow its rail. Pod size, rail count and port speeds come from your reference design.

The rails

Compute east-west

One port per accelerator, and every port on a different leaf. Eight rails across the pod, so a leaf failure costs one rail and never a whole node.

Eight of thirty-two nodes shown. Hover or tap a leaf to follow its rail. Pod size, rail count and port speeds come from the reference design being built.

The rail design is why the port maps matter so much on site. A lead in the wrong leaf port still lights, still passes a basic continuity check and quietly breaks the rail. Every port is mapped to a node and a rail before the first lead is made up, and the map is checked again at test.

04 / Speed steps100G to 1.6T on one physical layer

800G is shipping in hyperscale halls today and 1.6T is next. The lane count is what drives the connector, so the trunking that goes in now is specified for the step after the one being bought.

Parallel singlemode optics carry a transmit and a receive fibre per lane. At 100G and 400G that is four lanes and eight fibres, which sit in a 12-fibre MPO with four positions unused. At 800G the lane count doubles to eight, and sixteen fibres need a 16-fibre connector. 1.6T doubles the lane rate rather than the lane count, so it runs on the same sixteen fibres.

Tool 01Speed step explorerLanes, fibres and the connector

The lane count doubles

800G

Eight lanes, sixteen fibres. This is the step that changes the connector, and the reason to put base-16 trunking in before it is needed.

Changes at this step
  • The optics
  • The patch leads
  • The trunk, if base-12 went in
Stays in place
  • Base-16 trunks and cassettes
  • Containment and pathway
  • Labelling and as-builts

Parallel singlemode shown, because that is what carries the lane count up the curve. Duplex wavelength-multiplexed variants exist at every step for longer reach. Reaches are the usual figures for each optic class, not a guarantee for a particular part.

Why base-16 goes in now

A hall cabled in base-12 for 400G has to re-pull its trunks for 800G, or run conversion modules that strand fibres and add mated pairs. Base-16 trunks and cassettes installed from the start make 800G and 1.6T a change of optics and patch leads. MPO-16 is keyed differently from MPO-12, so the two cannot be mated by accident.

05 / The link budgetEvery mated pair spends budget

At 800G the optic’s loss allowance is small, and it does not grow with the speed. Every connector in the channel spends part of it.

A link from a node in one pod to a spine in the main distribution can pass through six mated pairs: patch cord to cassette, through the trunk, through the distribution frames and out again. With standard-grade connectivity at 0.5 dB a pair, six pairs and 300 m of singlemode already exceed the 3.0 dB allowance of an eight-lane parallel optic. The same channel built from ultra-low-loss components passes with margin.

Tool 02Link budgetMated pairs and route length
Mated pairs6
Route length300 m
Per mated pair0.50 dB standard grade, 0.25 dB ultra-low-loss. Maxima, on the same basis.
FibreSinglemode at 0.35 dB per km, 0.10 dB over the route shown.
Allowance3.0 dB for an eight-lane parallel singlemode optic at 500 m. The 2 km parallel variant allows 4.0 dB.
Standard connectivity3.10 dB0.10 dB overOver budget
Ultra-low-loss1.60 dB1.40 dB of headroomPasses

Illustrative planning figures. Every link is designed against the published channel specification of the optic being fitted.

This is why connector grade and cleanliness are design decisions rather than installation details. Ultra-low-loss components buy the margin that a contaminated endface, an extra patch or a re-route would otherwise use up, and at 1.6T the allowance is no larger than it was at 800G.

06 / Pre-terminatedPre-terminated, because the density demands it

At hyperscale link counts, field termination is the slowest and least repeatable part of the job.

Trunks and cassettes arrive terminated and measured, polarity is a scheme rather than a decision taken at each link, and the hall can be re-lit at the next speed step by changing patch leads.

01

Terminated and tested off site

Trunks and cassettes are terminated and measured under factory conditions and arrive with a result per fibre. Consistent loss across thousands of links comes from a controlled process.

02

Polarity as a scheme

One polarity method is chosen for the hall and carried through every trunk, cassette and lead. Nobody works it out at the cabinet, and a lead taken from stores is the right lead.

03

Inspect and clean, every mate

Every endface is inspected and cleaned to IEC 61300-3-35 before it is mated, including the ones that came out of a sealed bag. Contamination is the largest single cause of avoidable link failures at these densities.

04

Re-light, rather than re-pull

When the hall steps up a speed, the optics and the patch leads change. Base-16 trunks, cassettes, pathway and labelling stay exactly where they are, and the hall stays in service while it happens.

Containment and pathway, stays Optic Cass Cass Optic Patch lead, changes Cassette Base-16 trunk, stays Cassette Patch lead, changes
Replaced at the next speed stepInstalled once and left aloneLabels at both ends, and the as-builts, stay as they are

07 / The deployment lineGoods-in to handover, in one order

The same ten steps in the same order on every cabinet in the hall.

If a crew has to work out what comes next, cabinet six hundred will not match cabinet one. A written line, with the evidence each step produces, is what keeps a hall consistent across crews and shifts, and it is what lets a supervisor see at a glance where every cabinet stands.

DiagramThe deployment lineTen steps, every cabinet

Step 10 of 10

Handover

As-builts, patching schedules, the asset register and the full test results handed back in the operator’s own format, walked through at the hold point and signed off.

Hover or tap a step to hold it.

08 / The cabinetThe decisions inside one cabinet

A loaded GPU cabinet fills on power and weight long before it fills on rack units.

The elevation below is an example, drawn to show how the decisions inside one cabinet fit together: where the fibre field sits, how the switching and out-of-band are placed, how heavy nodes are handled and dressed, and why free rack units are not spare capacity. On a real deployment the elevation is the operator’s, and the cabinet is built to it U for U.

DiagramExample cabinet elevationSelect a detail
481216202428323640444852 AB Example elevation, 52U frame. Not to scale.

Detail 04

Compute nodes

Heavy, deep and awkward. The handling plan matters as much as the elevation.

example frame
52U
used in this elevation
40U
blanked, not spare capacity
12U
feeds to two rack PDUs
A+B

An example elevation, drawn to show how the decisions fit together. Not to scale.

09 / The standardThe operator issues the standard

At this scale the operator owns the method. The installer’s job is to work inside it exactly, in every cabinet.

Six documents usually come from the operator, and each one has a matching discipline on the floor.

Issued by the operator

The rack elevation

Built U for U and checked against the drawing before the cabinet is signed off. Where the kit does not match the drawing, it is raised rather than absorbed.

Issued by the operator

The labelling scheme

Applied to every port, panel, lead and cabinet at both ends, and carried into the as-builts and the patching schedule without a local variation.

Issued by the operator

The routing and dressing spec

Followed to the letter, including the parts that only matter in year three: bundle sizes, tie spacing, radius control and which side a cord leaves on.

Issued by the operator

The inspection and test plan

Worked to as written, with the hold points held. Nothing moves past one because the programme is tight, and the witness is invited rather than assumed.

Issued by the operator

The evidence requirement

Photographs at the operator’s checkpoints, against the cabinet reference, uploaded from the floor rather than assembled afterwards.

Issued by the operator

The asset register format

Populated in the operator’s schema, with their field names and validation, so the file imports on the day it lands instead of coming back for a re-key.

Where there is no standard

The installer should propose one in writing, have it approved before the first cabinet, and hold it as the standard for the rest of the hall. Writing it down once is cheaper than discovering, at cabinet three hundred, that two crews have been dressing the same leads two different ways.

10 / Waves and crewsThe hall fills in waves

A hall is released in waves of pods, and the crew ramps up and down with them.

Each wave runs the same sequence: cabinet build, node install, fabric cabling, then test and evidence. Waves overlap, so in a typical week one wave is being cabled while the next is being built and the one before is under test. The crew size follows the overlap, and the programme is worked back from the date the first pods have to be live.

Tool 03Wave programmeDrag the week
Engineers on site
Week 9
Drag the week

Week9of 18

Done

2 pods certified and handed over

Under way

Wave 2, fabric cabling. Wave 3, cabinet build.

Next

Wave 3, node install, week 10

pods handed over
2
engineers on site this week
34

An illustrative programme, drawn to show the shape of a two-pod wave, not a quote. Real wave sizes come from the release dates and the access the site can give.

11 / ScaleSizing the hall

The numbers that drive a deployment come from a small model: cabinets, nodes per cabinet, and the fabrics each node connects to.

Move the sliders to size a hall. Every figure below is built from the model shown, so you can see what drives it.

Tool 04Hall scalerCabinets and nodes per cabinet
Compute cabinets128
Nodes per cabinet4

Each block is a pod of eight cabinets128 cabinets, 16 pods

compute nodes
512
ports on the compute fabric
4,096
fibre links, every fabric
6,400
copper links, out-of-band
896
engineer-days to build, cable and dress
259
days to certify 100% of the links
9

Model8 cabinets per pod8 accelerators per node2 storage links per node10 leaves per pod, 8 uplinks each1 out-of-band link per node, 3 per cabinet160 links per engineer-day3 nodes and 3 cabinets per engineer-day900 links per test-day

Illustrative planning figures. Every deployment is surveyed and scoped before a programme or a price is committed.

Two things stand out at any size. Fibre links outnumber nodes many times over, because every accelerator has its own port on the compute fabric; and certification, at around 900 links a test-day in this model, is a real line in the programme rather than a final afternoon.

12 / Test and evidence100% certified, link by link

At handover the evidence is the deliverable: a result for every link, a photograph set for every cabinet, and records that import on the day they land.

Fibre is tested for insertion loss and with an OTDR, with launch and receive cables fitted so the first and last connectors are measured. Copper out-of-band links are certified as permanent links to TIA-568 and ISO/IEC 11801. Nothing is sampled.

ExampleTest recordsFacsimile

MDA-A03-U12-P01 to DH1-B07-U40-P01

Singlemode, tested at 1310 nmLaunch and receive cables fittedEndfaces inspected to IEC 61300-3-35

Pass

DH1-B07-U45-P12 to DH1-B07-U18-P12

Permanent link, Cat6ALimit: TIA-568 and ISO/IEC 11801Calibrated field tester

Pass

Facsimiles, built to show the shape of the record rather than a real result.

How the links are tested

Every link, no sampling

  • Fibre: insertion loss and OTDR, with launch and receive cables fitted on every link, and every endface inspected and cleaned before it is mated.
  • Copper: channel and permanent link testing to TIA-568 and ISO/IEC 11801, on calibrated field testers.
  • Defects: reported, remediated and retested before the link is signed off.
The handover pack

What exists at handover

  • Test results, link by link, fibre and copper
  • Photographs per cabinet, at the operator’s checkpoints
  • As-built drawings and containment routes
  • The patching schedule, both ends of every link
  • The asset register, in the operator’s format
  • Every deviation raised, with the approval against it
Example label

DH1-B07-U42-P24toMDA-A03-U10-P24

DH1
Data hall
B07
Row and cabinet
U42
Panel position
P24
Port
Both ends
Same rule at the far end

13 / How Optronix helpsOne owner for the fabric and the kit

Structured cabling, containment, power distribution and white-space fit-out are our core business. We design, supply, install, certify and migrate across the full data centre lifecycle for hyperscale, colocation and enterprise clients, across the UK, EMEA and Asia-Pacific.

01

Our engineers are ours

Optronix employs its engineers directly. The people on the hall are on our payroll and under our supervision, working to one written method on every shift.

02

Current at this density

A large share of our current work is AI and GPU deployment, where fibre density has moved to 800G. This is the work we are doing now, not work we are preparing for.

03

Evidence as standard

Certified links, photographs, as-builts and a populated asset register are part of the work, in the operator’s format, and not a line item we come back for.

One scope, from goods-in to handover

  • 01

    Receipt and asset capture

    Deliveries checked against the manifest, damage recorded, and every serial, tag and position captured before a box leaves goods-in.

  • 02

    Cabinet and rail build

    Cabinets set out to the floor grid, bayed, levelled and bonded, then rails and cable management built to the issued elevation.

  • 03

    Compute and GPU nodes

    Heavy accelerated nodes handled with lifting aids and two people, seated on their rails, cabled clear of the airflow path and powered on in order.

  • 04

    Power to the rack

    A and B feeds landed into the rack PDUs, outlets mapped to the elevation, cords cut to length, retained and dressed to the correct side.

  • 05

    Pathways and containment

    Tray, basket and in-row managers installed and loaded to the design fill, with bend radius kept on every fibre route.

  • 06

    Fabric cabling

    Compute, storage, in-band and out-of-band networks, installed as pre-terminated structured links with polarity managed as one scheme.

  • 07

    Labelling and records

    Every port, panel and lead labelled at both ends to the operator’s scheme, and the asset register and patching schedule built as the work goes in.

  • 08

    Test, certify, evidence

    100% of links certified, results recorded link by link, photographs per cabinet, and the pack handed back in the operator’s format.

  1. 01

    Survey and plan

    The reference design, the pod count and the first live date are enough to start: the scope, the crew shape and an outline programme come back worked from that date.

  2. 02

    Mobilise and build

    Goods-in, asset capture, cabinet build and node install, by directly employed engineers with a supervisor on every shift.

  3. 03

    Cable, test, certify

    Pre-terminated fabric cabling to the port maps, every endface inspected and cleaned, and 100% of links certified.

  4. 04

    Evidence and handover

    Records in the operator’s format, issued wave by wave and walked through at the hold points.

Related services

The rest of the series

The physical layer is covered in depth in structured cabling; the power, cooling and fabric design of the halls these deployments fill in planning an AI/GPU-ready data hall; and the shell-to-live programme around them in the data centre fit-out.

Hyperscale deployment

Tell us the hall and the date.

The reference design, the pod count and the date the first pods have to be live are enough to start. We will come back with the scope, the crew shape and an outline programme worked back from that date.

Share this guide with your team. Share on LinkedIn

Sources and notes

  • TIA-942 and ISO/IEC 11801: data centre cabling spaces, the distribution hierarchy and media.
  • IEEE 802.3: 400GBASE-DR4 and 800GBASE-DR8 parallel singlemode optics (802.3bs, 802.3df) and their channel insertion loss allowances; 1.6T on 200G lanes (802.3dj).
  • IEC 61300-3-35: visual inspection of fibre optic connector endfaces.
  • TIA-568.3 and the IEC 61280-4 series: optical fibre cabling test methods, insertion loss and OTDR, with launch and receive cables.
  • TIA-568.2 and ISO/IEC 11801: balanced copper permanent link and channel testing.
  • IEC 61754-7 family: MPO connector interfaces, including the 12-fibre and 16-fibre variants and their keying.
  • General industry practice for rail-optimised GPU fabrics, pre-terminated base-16 cabling and wave-based hall deployment.

General guidance, not a design specification. Loss figures, reaches, lane rates, the wave programme and the planning model are illustrative and depend on the specific optics, components and site. Confirm against the published specifications and an engineered design.