| title | AI-Native Data Center Pod - 10 MW Reference Architecture (PoC) |
|---|---|
| author | Antreas Christofi |
| date | 2026-02-11 |
| toc | true |
| toc-depth | 3 |
- BoD is the single normative source of requirements and evidence gates.
- Appendices are the single source of derivations and worked examples.
- Reference Architecture instantiates a topology that satisfies the BoD by reference only.
- Witness excerpt is a pointer and reconciliation summary; the authoritative witness lives in Appendix F.
Physics-First, Vendor-Agnostic
Purpose:
Define a facility capable of sustaining ~10 MW IT load for large-scale AI training with liquid-cooled racks, while remaining safe, observable, and operable across its lifecycle.
What this document is
- A set of engineering contracts between IT and facility
- A definition of envelopes (power, heat, water, telemetry)
- A framework for commissioning and evidence
What this document is not
- A vendor selection
- A construction drawing pack
- A detailed controls implementation
System Chain (intent)
Workload → Electrical Power → Heat Capture → Fluid Transport → Heat Rejection → Monitoring → Operations
Key Envelopes
| Domain | Target |
|---|---|
| IT load | 10 MW nominal |
| Rack cooling | Direct liquid to chip + CDU |
|
|
8–12 K initial assumption |
| Redundancy | N+1 at critical layers |
| Observability | Second-level telemetry |
- Electrical & mechanical engineering
- Controls & BMS teams
- Safety/compliance
- Operations
- IT boundary: rack manifolds and CDU interface
- Facility boundary: power delivery, fluid loops, heat rejection, controls
- Responsibility split defined at CDU inlet/outlet and rack PDU
- FW – facility water loop
- TC – technical coolant to racks
- CDU – coolant distribution unit
- HIL – hardware-in-loop testing
- KPI – sustainability & performance metric
The facility shall:
- Deliver up to 10 MW IT power with defined ramp limits
- Support dynamic AI workloads with bounded inrush
- Provide:
- N+1 at MV/LV transformation
- Selective coordination
- Power quality monitoring
Evidence
- FAT/SAT under staged load banks
- Ramp tests 0→100% and 100→0%
- Harmonics and THD reports
- Capture ≥ 95% of IT heat via liquid path
- Maintain rack inlet temperature within OEM envelope
- Provide:
- Dual-loop separation (FW / TC)
- Leak containment and detection
- Minimum
$\Delta T$ target 8 K
Evidence
- Calibrated heat balance
- Flow vs load linearity
- CDU failover test
- Water quality per OEM specification
- Treatment for corrosion/biological control
- Leak detection at:
- rack
- CDU
- header
Evidence
- Lab analysis
- Pressure decay tests
- Containment verification
- All critical paths instrumented
- Time-synchronized telemetry
- Alarms classified:
- Safety
- Service
- Advisory
Evidence
- Telemetry acceptance
- Alarm storm testing
- Loss-of-signal scenarios
- Measure PUE, WUE, CUE boundaries
- Reportable metering at:
- utility entry
- IT PDUs
- cooling plant
Evidence
- KPI reconciliation
- Sensor calibration records
- Utility → MV → LV
- Static transfer where required
- Rack PDUs with branch monitoring
- Grounding and bonding strategy
- FW loop to dry coolers
- CDU layer with TC distribution
- Residual air path for ancillaries
- BMS + DCIM integration
- State machines for:
- ramp
- fault isolation
- maintenance modes
- EPO logic
- Leak isolation
- Fire strategy compatible with liquid cooling
This section records the key architectural decisions made within the admissible physical envelope established in Appendices A–F.
The appendices prove feasibility.
This section explains selection.
Each decision represents a chosen operating point within a broader feasible region, balancing thermodynamic efficiency, reliability, operability, and deployment constraints.
Decision
Warm-water direct-to-chip liquid cooling with CDU-mediated separation between facility water (FW) and technical coolant (TC).
Alternatives Considered
- Air cooling
- Rear-door heat exchangers (RDHx)
- Full immersion cooling
- Hybrid air/liquid approaches
Tradeoffs
| Approach | Advantages | Disadvantages |
|---|---|---|
| Air cooling | Simpler infrastructure, mature operational model | Cannot sustain ≥100 kW/rack without excessive airflow, fan power, and acoustic burden |
| RDHx | Compatible with existing air-cooled hardware | Retains airflow dependency and heat rejection inefficiencies |
| Immersion | Excellent heat transfer and density potential | Operational complexity, hardware compatibility constraints, vendor coupling |
| Direct-to-chip (selected) | Highest heat capture efficiency, scalable to ≥150 kW/rack, vendor-neutral facility interface | Requires liquid infrastructure and leak containment |
Rationale
Direct-to-chip liquid cooling maximizes primary heat capture efficiency while preserving hardware and operational flexibility.
It aligns with the ≥95% liquid heat capture contract and enables rack densities beyond the practical limits of air cooling.
The CDU separation preserves facility independence from IT coolant chemistry and hardware refresh cycles.
Sensitivity
If rack densities remained below approximately 40 kW/rack, air or hybrid cooling approaches could remain viable.
If immersion becomes operationally standardized across vendors, immersion may become favorable in future revisions.
Decision
Design liquid temperature rise of approximately 12 °C (envelope 8–15 °C).
Alternatives Considered
- Low
$\Delta T$ (5–8 °C) - High
$\Delta T$ (15–20 °C)
Tradeoffs
| Advantages | Disadvantages | |
|---|---|---|
| Lower | Lower component temperature gradients, simpler control | Higher required flow rates, larger pipes, higher pumping power |
| Higher | Lower flow rates, smaller pipes, reduced pumping power | Reduced thermal control margin, potential component temperature risk |
| 12 °C (selected) | Balanced flow rate, pumping power, and thermal margin | Requires coordinated control |
Rationale
The selected
This choice reduces pipe diameter, pump sizing, and parasitic energy without compromising component safety.
Sensitivity
Lower
Higher
Decision
Maximum mechanical fault domain size of approximately 1 MW.
Alternatives Considered
- Smaller zones (250–500 kW)
- Larger zones (2 MW or greater)
Tradeoffs
| Zone Size | Advantages | Disadvantages |
|---|---|---|
| Smaller | Greater fault containment | Increased infrastructure complexity and cost |
| Larger | Reduced infrastructure overhead | Larger thermal impact during failures |
| 1 MW (selected) | Balanced containment, manageable infrastructure complexity | Moderate additional infrastructure overhead |
Rationale
A 1 MW zone size ensures that hydraulic or cooling failures remain locally recoverable without propagating to the full facility.
This aligns mechanical fault domains with operational recoverability and commissioning isolation.
Sensitivity
Higher availability targets may justify smaller zones.
Lower capital cost priorities may justify larger zones.
Decision
Electrical distribution blocks of approximately 2 MW.
Alternatives Considered
- Smaller blocks (≤1 MW)
- Larger blocks (≥5 MW)
Tradeoffs
| Block Size | Advantages | Disadvantages |
|---|---|---|
| Smaller | Better fault isolation | Increased electrical infrastructure complexity |
| Larger | Reduced infrastructure overhead | Larger blast radius during faults |
| 2 MW (selected) | Balanced coordination, scalability, and fault containment | Moderate infrastructure overhead |
Rationale
2 MW electrical blocks align with mechanical zoning while maintaining manageable switchgear and protection complexity.
This structure supports selective coordination and staged energization during commissioning.
Sensitivity
Higher grid instability environments may favor smaller blocks.
Highly stable utility environments may support larger blocks.
Decision
Primary reliance on dry cooling, with optional heat reuse as a secondary path.
Alternatives Considered
- Mandatory heat reuse
- Evaporative cooling primary systems
Tradeoffs
| Strategy | Advantages | Disadvantages |
|---|---|---|
| Mandatory reuse | Maximum energy recovery | Introduces dependency on external heat consumers |
| Evaporative primary | High efficiency | High water consumption |
| Dry primary, reuse secondary (selected) | Independence, water conservation, operational robustness | Slightly lower peak thermal efficiency |
Rationale
The facility must remain fully operable independent of external heat consumers.
Heat reuse is supported but must not introduce operational dependency.
Sensitivity
Sites with guaranteed heat offtake may increase reuse integration.
Decision
Digital twin operates in bounded supervisory mode, without authority over safety interlocks.
Alternatives Considered
- Full autonomous control
- Monitoring-only model
Tradeoffs
| Mode | Advantages | Disadvantages |
|---|---|---|
| Autonomous | Maximum automation potential | Increased safety and validation risk |
| Monitoring only | Maximum safety | Limited operational optimization |
| Supervisory bounded (selected) | Balanced safety and optimization | Requires defined authority boundaries |
Rationale
The twin provides operational optimization while preserving deterministic safety behavior.
Safety-critical interlocks remain hardware- or PLC-enforced.
Sensitivity
Future validated models may expand operational authority.
Decision
Facility cooling architecture remains vendor-agnostic at the CDU boundary.
Alternatives Considered
- Hardware-specific cooling integration
Tradeoffs
| Approach | Advantages | Disadvantages |
|---|---|---|
| Vendor-specific | Maximum optimization for specific hardware | Reduces flexibility, increases lock-in |
| Vendor-agnostic (selected) | Maximum flexibility, supports hardware evolution | Requires interface abstraction |
Rationale
Facility lifetime exceeds IT hardware lifetime.
Vendor-neutral interfaces preserve long-term adaptability.
The selected architecture represents a balanced operating point within the physically admissible region defined by Appendices A–F.
The chosen decisions prioritize:
- Fault containment
- Operational independence
- Thermal and hydraulic efficiency
- Commissionability
- Long-term adaptability
These selections establish a reference architecture that is physically grounded, operationally robust, and adaptable to evolving AI workloads.
| Item | Value (initial) |
|---|---|
| Heat to liquid | 9.5–10 MW |
|
|
10 K |
| Estimated flow | derived in Appendix B |
| CDU capacity | N+1 |
| Item | Value |
|---|---|
| IT power | 10 MW |
| Ancillary | ~8–12% |
| Diversity | workload dependent |
- Header sizing per hydraulic model
- Treatment plant sized for volume turnover
- 1 s resolution for critical paths
- 10 s for secondary
- Retention ≥ 24 months
For each contract:
- Method
- FAT
- SAT
- IST
- HIL
- Instrumentation
- Calibrated meters
- Flow/temperature pairs
- Power analyzers
- Acceptance
- Numeric criteria
- Stability windows
- Negative tests
- Artifacts
- Reports
- Traces
- As-built configs
- Change governance
- Maintenance schedules
- Alarm philosophy
- Training requirements
- Failure drills
- Version: Draft 1.0
- Nature: Engineering Basis of Design
- Ownership: Multidisciplinary
Appendix Index (authoritative derivations and evidence mapping)
- Appendix A — Workload → Boundary Projection
- Appendix B — Thermal Physics and Digital Twin Basis
- Appendix C — Hydraulics, Piping, and
$\Delta P$ Budgets - Appendix D — Sustainability Model and KPI Closure
- Appendix E — Cross-Domain FMEA and Commissioning Mapping
- Appendix F — CausalCompute Alignment (Witness Workload Closure)
This appendix translates arbitrary AI training objectives into the electrical and thermal quantities visible at the facility boundary.
The facility is intentionally workload-agnostic.
Therefore the mathematics is used to generate:
- admissibility envelopes, not a single model instance
- power transients, not GPU counts
- boundary quantities (
$P_{IT}$ ,$\Delta P$ ,$dP/dt$ )
| Symbol | Meaning | Units |
|---|---|---|
| model parameters | params | |
| total training tokens | tokens | |
| wall-clock budget | s | |
| FLOPs per param-token (dense ≈ 6) | FLOPs/(param·token) | |
| sustained training efficiency | – | |
| compute efficiency at IT boundary | FLOPs/s/W | |
| IT electrical power | W | |
| step change in power | W | |
| ramp rate | W/s |
Total algorithmic work:
Required sustained throughput:
Mapping to electrical boundary:
Admissibility condition
For this pod:
This relation does not size the facility to one model;
it checks whether any proposed workload lies inside the envelope.
AI training creates power transients through:
- job start/stop
- collective synchronizations
- checkpoint I/O
- failure recovery
Detailed behavior is reduced to enforceable limits:
Recovery objective:
The facility admits any workload whose boundary image lies in:
Workload mathematics therefore act as the projection operator into this admissible set.
Locked design
Agnostic design
-
Primary – Scheduler Policy
- staged energization
- anti-herd starts
- block-local ramp sequencing
-
Secondary – Power Capping
- per-rack limits
- collective smoothing
-
Tertiary – BESS/UPS
- grid event absorption
- residual transient buffer
Acceptance requires replay of representative traces:
- voltage sag within limits
- UPS not overloaded
- CDU controls stable
-
$T_{out}$ within envelope
In the agnostic formulation, workload mathematics serve to map arbitrary ML objectives into boundary quantities rather than prescribe a specific model instance.
This appendix produces:
- IT power envelope (
$P_{IT}$ ) - transient limits (
$\Delta P$ ,$dP/dt$ ) - inputs to thermal sizing (Appendix B)
- acceptance tests for commissioning
A single worked witness is maintained in Appendix F (CausalCompute alignment).
All numeric witness values used by Appendices B–E are sourced from Appendix F.
End of Appendix A
This appendix closes the thermodynamic and control loop for the single witness workload defined in Appendix F:
$P_{IT} \approx 4.77\ \text{MW}$ -
$\dot Q \approx P_{IT}$ (all IT power becomes heat, first order) - Design liquid rise:
$\Delta T = 12^\circ\text{C}$ (envelope 8–15 °C) - Mechanical zoning:
$1\ \text{MW}$ maximum fault domain per zone
It produces:
- Required coolant flow
$\dot m,\ \dot V$ at the pod boundary - Per-zone flow targets
- Pumping power order-of-magnitude from
$\Delta P$ budgets - A control-oriented thermal state model and time constants
- Commissioning identification and validation gates for the twin
| Symbol | Meaning | Units |
|---|---|---|
| heat rate | W | |
| IT electrical power | W | |
| mass flow rate | kg/s | |
| volumetric flow rate | m$^3$/s | |
| fluid density (water, 25–30 °C) | kg/m$^3$ | |
| specific heat (water) | J/(kg·K) | |
| temperature rise | K or °C | |
| pressure differential | Pa | |
| pump + motor + VFD efficiency | – | |
| effective thermal capacitance | J/K | |
| thermal time constant | s |
Water properties (PoC design basis)
$\rho \approx 997\ \text{kg/m}^3$ $c_p \approx 4180\ \text{J/(kg·K)}$
For an incompressible single-phase coolant loop:
Rearranging gives required mass flow:
Using density to convert to volumetric flow:
With
Compute denominator:
So:
Pod-level result (witness case)
For a zone cap of
Per-zone result
Witness load is served by approximately five 1 MW zones in aggregate:
Nominal flow if distributed across 5 zones:
This matches the pod-level 95 L/s within rounding and distribution overheads.
For a 100 kW rack:
Assume a PoC zone budget:
$\Delta P_{zone} = 100\ \text{kPa} = 100{,}000\ \text{Pa}$ $\dot V_{zone} \approx 0.020\ \text{m}^3/\text{s}$ $\eta_{pump} = 0.70$
If the witness workload uses ~5 zones:
This is a control-oriented abstraction sufficient for estimation, anomaly detection, and supervisory setpoint control.
At steady state:
Dominant time constant:
Illustrative inventory:
$V_{eq} \approx 5\ \text{m}^3$ $m_{eq} \approx 5000\ \text{kg}$
With
For sampling interval
\dot m[k],c_p,(\hat T_{out}[k]-T_{in}[k]) \right) $$
Telemetry sources:
-
$\dot Q[k]$ : rack/block power metering -
$\dot m[k]$ : flow meter -
$T_{in}[k]$ : supply temperature -
$T_{out}[k]$ : return temperature probe
Procedure:
- Hold
$T_{in}$ approximately constant - Hold
$\dot m$ constant - Apply a bounded step or ramp in
$\dot Q$ - Measure
$T_{out}(t)$ - Fit
$\hat C$ by minimizing prediction error
Acceptance (PoC intent):
-
Pod-level flow: $$ \dot V \approx 0.095\ \text{m}^3/\text{s} \approx 95\ \text{L/s} $$
-
Zone-level flow: $$ \dot V_{zone} \approx 0.020\ \text{m}^3/\text{s} \approx 20\ \text{L/s} $$
-
Rack flow: $$ \dot V_{rack} \approx 2\ \text{L/s} $$
-
Zone time constant (illustrative): $$ \tau \approx 250\ \text{s} $$
This appendix converts the thermal requirements from Appendix B into buildable hydraulic quantities:
- header diameters from the single witness flow (
$\approx 95\ \text{L/s}$ ) - per-zone distribution at
$\approx 20\ \text{L/s}$ - velocity and
$\Delta P$ budgets - pumping power and sensitivity
- fault-domain containment rules
All calculations are SI; water at 25–30 °C is assumed.
For design velocity
From Appendix B:
Use
Interpretation
- Primary header class ≈ DN250 (DN250–DN300 depending on allowances)
- The full 10 MW pod (200 L/s) would approach DN350 class
Zone flow:
Using
Practical class: DN100–DN125 per 1 MW zone.
For 100 kW rack:
If internal manifold velocity target
Approx: DN40 class at rack interface.
| Segment | Target |
|---|---|
| Rack cold plates & hoses | ≤ 40 kPa |
| CDU HEX + internals | ≤ 40 kPa |
| Distribution & valves | ≤ 20 kPa |
| Total zone | ≤ 100 kPa |
Sensitivity:
| Location | Limit | Reason |
|---|---|---|
| Primary headers | ≤ 2.0 m/s | erosion, noise, transients |
| Secondary runs | ≤ 1.5 m/s | controllability |
| Rack hoses | ≤ 1.2 m/s | connector wear |
| Drain lines | ≥ 0.6 m/s | sediment transport |
Rapid valve closure can generate:
Mitigations:
- rate-limited actuators
- surge arrestors
- controlled pump ramps
- commissioning verification of worst-case closures
Containment rule:
A single failure shall not exceed 1 MW thermal radius.
Therefore:
- isolation valves at each zone boundary
- no common headers without sectionalization
- wet components segregated from electrical rooms
Design spill volume per zone:
With 20 L/s and 10 s detection:
Minimum per zone:
- supply
$T_{TC,in}$ - return
$T_{TC,out}$ - flow
$\dot m$ or$\Delta P$ proxy - valve position
- leak detection (cable + point)
Acceptance tests derived from this appendix:
-
Flow verification: $$ \dot V_{zone} \ge 20\ \text{L/s} $$
-
$\Delta P$ envelope: $$ \Delta P_{zone} \le 100\ \text{kPa} $$ -
Velocity audit: confirm
$v \le 2\ \text{m/s}$ in primaries -
Transient test: staged valve closure without excursions
-
Leak isolation: single-zone containment proven
This appendix converts the single witness workload into auditable sustainability quantities:
- annual energy (IT and facility)
- heat export potential (
$\text{MWh}_{th}$ ) - water sensitivity under adiabatic/evaporative augmentation
- KPI boundaries and required metering
- a PoC measurement plan that makes reporting defensible
The intent is not “green claims,” but metered contracts.
From Appendix F:
Design intent:
$\text{PUE}_{design} = 1.15$ - utilization factor
$U = 0.70$ (example)
Facility power:
Compute:
Thermal energy available at IT boundary (first order):
Annual thermal energy:
Heat pump relation (if needed):
If
Latent heat of vaporization:
Evaporation mass flow to reject heat
For
Daily water (if operating adiabatic continuously):
Annual at utilization
Requirements:
- computed as a time series with time-aligned meters
- reported as distributions/percentiles, not a single number
Where make-up water exists:
Interval form:
Exported thermal power:
Exported energy:
If grid intensity is a time series
Electrical:
- Utility import meter (true RMS, time-stamped)
- Plant distribution meters (cooling, UPS losses, BESS auxiliaries)
- IT boundary meters per block / PDU / RPP (depending on design)
Thermal / Water:
- Heat meter at export interface:
$\dot V + T_{supply} + T_{return}$ - FW/TC loop meters per zone:
$\dot V$ ,$T_{in}$ ,$T_{out}$ ,$\Delta P$ - Make-up water meter where adiabatic exists
Time synchronization:
- PTP/NTP such that meter timestamps are aligned and drift is monitored
Heat export must not become a single point of failure.
Contractual rule:
- Export path must fail open to ambient rejection.
- Off-taker unavailability shall not cause derate of IT load.
This implies bypass and buffering at the export interface.
Using
| Quantity | Result |
|---|---|
| IT annual energy |
~29,300 MWh/yr |
| Facility annual energy |
~33,600 MWh/yr |
| Overhead energy | ~4,300 MWh/yr |
| Heat potential |
~29,300 |
| Evaporation rate (adiabatic continuous) | ~1.95 L/s |
| Daily water (adiabatic continuous) | ~168 |
| Annual water (adiabatic continuous, |
~43,000 |
Interpretation: even at ~5 MW class loads, adiabatic hours must be explicitly controlled and metered.
To make the PoC defensible:
- Verify
$P_{IT}(t)$ ,$P_{fac}(t)$ and compute$\text{PUE}(t)$ - Verify heat balance closure: $$ P_{IT} \approx \rho c_p \dot V \Delta T + \text{loss terms} $$
- If export exists, verify heat meter accuracy and availability
- If adiabatic exists, verify make-up meter and compute
$\text{WUE}(t)$ - Time-sync audit: confirm timestamp alignment across power + flow + temperature
Acceptance intent: all reported KPIs must be reproducible from raw time-series data with documented calibrations.
This appendix applies a structured Failure Modes and Effects Analysis (FMEA) to the
10 MW PoC architecture while explicitly referencing the single witness workload (Appendix F):
$P_{IT} \approx 4.77\ \text{MW}$ $\dot V \approx 95\ \text{L/s}$ $\dot V_{zone} \approx 20\ \text{L/s}$ -
$\tau \approx 250\ \text{s}$ (illustrative)
The objective is to:
- Identify failure modes that could violate the envelopes derived in Appendices A–D
- Quantify risk using the common O/S/D scale
- Map each high-RPN item to a commissioning evidence gate
Risk Priority Number:
| Range | Action |
|---|---|
| ≥300 | Architectural redesign |
| 200–299 | Mandatory mitigation + tests |
| 120–199 | Mitigation required |
| 60–119 | Manage with monitoring |
| <60 | Accept / monitor |
| Failure Mode | O | S | D | RPN | Witness Impact | Mitigation | Commissioning Evidence |
|---|---|---|---|---|---|---|---|
| Block ramp exceeds envelope | 5 | 8 | 4 | 160 | Violates |
Scheduler staging + capping | HIL replay of ramps |
| Protection mis-coordination | 3 | 9 | 6 | 162 | Upstream trip near block limits | Settings governance | Selectivity test |
| UPS bypass failure | 2 | 9 | 6 | 108 | Loss of power quality | Periodic transfer tests | Bypass SAT |
| Harmonics at PCC | 4 | 5 | 5 | 100 | PQ breach | Filters + monitoring | PQ report |
| Powered but uncooled | 3 | 10 | 6 | 180 | Thermal runaway | Hard interlocks | Interlock FAT |
| Failure Mode | O | S | D | RPN | Witness Link | Mitigation | Evidence |
|---|---|---|---|---|---|---|---|
| Loss of flow in 1 MW zone | 4 | 9 | 5 | 180 |
|
Dual CDU + alarms | Flow fail test |
| Valve hunting | 5 | 6 | 6 | 180 | Violates |
Rate limits | Step response |
| Filter fouling | 6 | 5 | 4 | 120 | DP trending | PM records | |
| Leak at rack | 4 | 8 | 4 | 128 | ~200 L spill | Trays + isolation | Leak drill |
| Air ingress | 4 | 7 | 6 | 168 | Oscillatory |
Degassing | Stability test |
| Failure Mode | O | S | D | RPN | Witness Link | Mitigation | Evidence |
|---|---|---|---|---|---|---|---|
| Sensor drift | 6 | 7 | 6 | 252 | Bias in |
Redundancy | Calibration |
| Model obsolete | 5 | 7 | 7 | 245 | Wrong |
BIM versioning | Re-fit gate |
| Time sync loss | 5 | 6 | 7 | 210 | KPI corruption | PTP/NTP | Audit |
| Unsafe actuation | 3 | 10 | 6 | 180 | Flow violation | Guardrails | Negative test |
| Failure Mode | O | S | D | RPN | Witness Link | Mitigation | Evidence |
|---|---|---|---|---|---|---|---|
| Unmetered PUE | 6 | 6 | 7 | 252 | KPI invalid | Meter plan | Data audit |
| Stranded heat | 6 | 6 | 6 | 216 | Export not utilized | Fail-open | Contract |
| Water cap breach | 4 | 8 | 6 | 192 | Adiabatic overuse | Dry baseline | WUE logs |
- Replay ramp to confirm: $$ \left|\frac{dP}{dt}\right| \le 0.1\ \text{MW/s} $$
- Verify per-block metering and coordination
- Bypass and transfer tests
- Prove: $$ \dot V_{zone} \ge 20\ \text{L/s} $$
- Confirm: $$ \Delta P_{zone} \le 100\ \text{kPa} $$
- Leak isolation within single zone
- RMSE: $$ \le 0.5^\circ\text{C} $$
- Fault injection for drift and flow loss
- Reproduce
$\text{PUE}(t)$ from raw meters - Verify heat meter closure
- Sensor integrity for twin (
$RPN=252$ ) - KPI auditability (
$RPN=252$ ) - Model validity after changes (
$RPN=245$ )
Mitigation requires process + engineering, not hardware alone.
The witness workload demonstrates:
- envelopes can be met
- failure modes are observable
- each risk maps to a testable gate
No
This appendix demonstrates that the envelopes defined in the Basis of Design can host a representative AI training workload using an external first-principles sizing engine (CausalCompute, Steps 0–2).
This appendix is:
- a traceability bridge between ML intent and facility physics,
- an evidence layer showing that the BoD envelopes are internally consistent,
- a generator of measurable PoC acceptance criteria.
This appendix is not:
- a redesign of the facility,
- a procurement specification,
- a network topology design,
- a replacement for the thermo-hydraulic derivations in Appendices A–C.
The BoD is a pre-production engineering contract, not an as-built system.
It defines envelopes and evidence gates that detailed design must later satisfy.
CausalCompute is used here as a PoC witness engine.
Allowed conclusions:
- envelope admissibility at the facility boundary (MW, L/s, kW/rack)
- sensitivity to explicit assumptions (
$\eta$ ,$\Delta T$ , checkpoint policy)
Disallowed conclusions:
- procurement GPU counts or guarantees
- network topology or congestion claims
- ramp compliance verification
- reliability closure (N+1) or operational readiness
BoD — Bottom-Up (Facility-First)
Power → Heat → Fluid → Space → Operations
|________________________________
|
Workload → FLOPs → Time → GPUs → Power → Heat
CausalCompute — Top-Down (Workload-First)
Convergence: heat is the handoff point.
The tool delivers an implied heat rate:
The BoD accepts or rejects that heat rate using its
| Domain | Primary Reference |
|---|---|
| Workload → boundary mapping | Appendix A |
| Thermal closure | Appendix B |
| Hydraulics & |
Appendix C |
| Sustainability accounting | Appendix D |
| FMEA & commissioning | Appendix E |
| Item | Value |
|---|---|
| Parameters | |
| Tokens | |
| Deadline | 60 days |
| FLOPs/param-token | 6 |
| Item | Value |
|---|---|
| Device sustained | 0.875 PFLOP/s |
| Device memory | 80 GB |
| 0.30 | |
| 0.80 |
| Envelope | Value |
|---|---|
| IT power | ≤ 10 MW |
|
|
12 °C (8–15) |
| Rack cap | 100 kW |
| Zone cap | 1 MW |
| Flow rule | ~20 L/s per MW |
| Quantity | Result |
|---|---|
| Required throughput | 1.41 EFLOP/s |
| Token rate |
|
| Checkpoint size | 810 GB |
| Max step time | 69 s |
| Item | Result |
|---|---|
| GPUs | 6,208 |
| Nodes | 776 (8 GPU/node) |
| Parallelism | DP/TP/PP = 97/16/4 |
| Step time | 68.9 s ≤ 69 s |
| Memory/device | 52 GB ≤ 80 GB |
| Quantity | Result | Envelope | Status |
|---|---|---|---|
| IT power | 4.77 MW | ≤ 10 MW | Aligned |
| Facility power | 5.49 MW | — | — |
| PUE | 1.15 | BoD target | aligned |
| Quantity | Result (SI) | Practical Equivalent | Rule | Status |
|---|---|---|---|---|
| Heat | 4.77 MW | — | 9.5–10 MW cap | Aligned |
| Coolant flow | 0.095 m$^3$/s | 95 L/s (5,700 L/min) | 20 L/s per MW | Aligned |
| Per-zone load | ~0.95 MW | — | ≤ 1 MW | Aligned |
Using Appendix B:
| Item | Result (SI) | Practical Equivalent | Limit | Status |
|---|---|---|---|---|
| Racks | 49 | — | — | — |
| IT per rack | 97 kW | (≈ 331,000 BTU/h) | 100 kW | Aligned |
- Flow per zone ≈ 20 L/s (≈ 1,200 L/min)
- Header class consistent with DN200–DN250
- Pumping power remains kW-scale
The witness workload lies inside all BoD envelopes for electrical, thermal, hydraulic, and rack constraints.
| Change | Effect |
|---|---|
| GPUs ↑ → IT power ↑ | |
|
|
Flow ↑ ~50% |
| Shorter checkpoint | Storage BW ↑ |
| Step time ↑ → GPUs ↑ |
-
Power → heat closure: $$ \dot Q \approx P_{IT} $$
-
Flow law: $$ \dot m = \frac{\dot Q}{c_p , \Delta T} $$
-
Ramp compliance: per Appendix A
-
Digital twin accuracy: per Appendix B
End of Appendix F
This document instantiates the Engineering Basis of Design (BoD) into a coherent 10 MW reference topology.
This document contains no original requirements and no original derivations.
Requirements are defined in sections above.
This document defines a 10 MW AI-native data-center pod derived from physics rather than historical templates:
Workload → FLOP → Watts → Heat → Water → Space → Operations
Key principles:
- Bounded fault domains — 2 MW electrical blocks aligned to 1 MW mechanical zones prevent local failures from becoming site events.
- Explicit transients — sized against verified ramp envelopes (
$\Delta P$ ,$dP/dt$ ), not steady-state MW alone. (See Appendix A.) - Closed-loop operations — a physics-based digital twin supervises within hard interlocks and auditable safety limits. (See Appendix B.)
The architecture is intentionally vendor-neutral and deployable across European jurisdictions.
AI workloads have broken classical datacenter assumptions:
- Power density: 5–15 kW/rack → 30–150+ kW/rack
- Heat path: air → liquid-dominated
- Utilization: workload-bound, not facility-bound
- Deadlines: training schedules compete with time-to-power
Therefore infrastructure must be derived from workload physics, not legacy layouts.
- Grid capacity and ramp constraints
- European limits on water, noise, and heat rejection
- Reliability targets (N+1 / optional 2N)
- Mixed customer hardware maintainability
- Cooling topology (D2C/RDHx/hybrids)
- Electrical architecture (UPS/BESS placement)
- Spatial modularization
- Digital-twin control strategy
- Model size, tokens, deadline
- Implied sustained FLOP/s
- Availability & jitter tolerance
- Hardware class envelope (TDP,
$\Delta T$ limits)
Workload-to-boundary mapping is defined in Appendix A.
Witness closure via CausalCompute is defined in Appendix F.
- IT Load: 10 MW envelope (BoD)
- Rack baseline: 100 kW class
- Cooling: warm-water D2C with CDU layer
- Electrical blocks: 2 MW fault/coordination domains
- Mechanical zones: 1 MW maximum fault domains
Derived flow rules,
- Utility → MV → LV distribution aligned to 2 MW blocks
- Per-block metering at the IT boundary (BoD observability contract)
- Selective coordination implemented per BoD electrical contract
- Interlock: rack enable requires cooling availability (BoD thermal contract)
Transient contract is defined in Appendix A and verified per Appendix E.
- FW loop to dry coolers (baseline)
- CDU layer separating FW/TC
- Zone isolation at 1 MW boundaries
- Residual air path for ancillaries
Thermal closure and twin model basis are defined in Appendix B.
Hydraulics, pipe classes, and
- Metering topology and KPI computation per Appendix D
- Export interface is fail-open to ambient rejection per Appendix D
- KPI reporting is time-series, not scalar summaries
The twin uses the first-order energy balance defined in Appendix B, operating in modes:
Observe → Recommend → Bounded Actuate.
Authority boundaries: cannot override hard interlocks (BoD safety/thermal contracts).
- Progressive energization
- Negative testing
- Twin validation before authority
- Evidence-based acceptance
Test matrix and FMEA mapping are defined in Appendix E.
- Electrical studies (load-flow, protection, harmonics)
- Hydraulic model and transient analysis
- Controls I/O and interlock matrix
- Twin data schema
- BIM constraints
Infrastructure decisions are derivations from physics and workload, not brand templates.