Build, Returns & Procurement Plan
Drag-and-drop the racks (20 kW ceiling, live power) · full BOM with street pricing · selectable RAM (1 TB default per IT) and firewall with tenant scenarios · scored third-party picks · a returns model that separates RTX servers from DGX, splits shared infra, decays the rental rate, and shows a capital-recovery schedule · and a full build-out view.
▸ Drag any unit to reorder it, or drop it into the other rack. Or use the ↑ ↓ ⇄ buttons on each unit. Power gauges update live against the 20 kW per-rack ceiling. Hover a unit for details.
Per your note: no DGX Spark in Rack 1 — it's all RTX. Rack 1 holds 2 GPU servers + all shared infra; the 3rd RTX server sits in Rack 2.
The container manufacturer has now confirmed 20 kW per rack (not 15) when ambient stays below 35 °C or evaporative spray is on — so Rack 1 runs 2 servers + full infra at ~16 kW comfortably, and you can drag the 3rd server into Rack 1 to see it approach the ceiling. Rack 1 is a complete, sellable cluster on its own; Rack 2 is scaling you can defer. RAM is selectable below (1 TB / 24×48GB default per IT, full bandwidth). See the thermal answer below.
SHARED — serves the whole deploymentLOCAL — dedicated to this rack Blanking (auto-fills empty U)
Shared-infrastructure reach — per-rack or whole-deployment?
The compute-fabric gear (storage HA pair, 400G switch, head node, OOB switch) is one set for the whole deployment, physically in Rack 1; other racks' servers connect back over inter-rack 400G. Power gear (PDU/UPS/ATS) is per-rack.
Shared element
Scope
Supports ~
First to run out?
⚡Thermal & power reality check
Answering your two questions, with the responses you got back from local IT and from the container manufacturer, plus what the electrical schematic shows.
1 · Can a single container rack actually sustain ~20 kW (cold-aisle supply + hot-aisle containment), or is 15 kW a hard ceiling?
~20 kW per rack is supported — 15 kW is not a hard ceiling. The manufacturer confirms each rack can take up to 20 kW (≈160 kW for the whole container), conditioned on keeping the local environment below 35 °C, or using evaporative spray to pull inlet air below 35 °C. The cooling progresses air → evaporative → AC, so the binding variable is outside ambient temperature and humidity, not the rack itself.
Manufacturer: "Each rack can support up to 20 kW (160 kW for one container) — provided the operating environment stays below 35 °C, or effective spraying lowers inlet air below 35 °C."
Practical effect: this build runs 2 servers + full infra at 15.75 kW in Rack 1 (comfortable), and a rack can physically hold 3× 6.5 kW servers = 19.5 kW if you ever go pure-compute.
2 · Each server is ~6.5 kW in 5U with high front-to-back CFM. Is that per-box heat density OK for the containment?
Yes — provided every unused rack opening is blanked. The manufacturer is explicit that to keep the high CFM moving through the servers you must install blanking panels in all unused openings so cold air can't bypass the servers straight to the rear. Those panels ship in the accessory kit (and are in this BOM).
Manufacturer: "Install blanking panels in any unused rack openings — this prevents cold air bypassing the servers to the rear. The blanking panels are included in the accessory kit. The 250 A breaker is dedicated to the servers; we have two, each feeding the PDUs on one side of the rack."
Local IT adds a sizing + de-risk view: with two 250 A breakers (one per PDU side, confirmed on the schematic you sent) and an 80% continuous-load rule, the container tops out around 18× 6.5 kW servers unless the gear/breakers are rated for 100% continuous — and they recommend starting at 6 GPU cards per server to validate thermals before populating all 8.
Local IT: "Max ~18× 6.5 kW servers to stay at 80% continuous (unless rated for 100%). We're not familiar with these containers — biggest factor is outside ambient + humidity. Going with 6 cards per server to start is a smart move until we validate thermals."
What this changes: the rack model above now uses a 20 kW ceiling. Per rack you can fit 2 servers + infra (15.75 kW) or 3 pure-compute servers (19.5 kW). Container-wide, plan for ~18 servers (conservative, 80% continuous) up to ~24 (the data switch / theoretical power limit) — that ~18–24 range is exactly the “max servers” band in the returns model. De-risk: consider a 6-card first article (~5.3 kW/server) for the POC, then populate to 8 once thermals are validated.
02Bill of materials + pricing
All pricing is indicative street/list. Each per-server line shows price position, lead time, and a faster alternative if it's >4 weeks out. Every part # is a buy link. RAM is selectable in the panel below (default 1 TB / 24×48GB per IT review — full bandwidth).
Per-server total $142,500 (street) — RAM is the swing variable: 768 GB (24×32GB) is the cheapest baseline, 1 TB (24×48GB) is the IT-recommended default selected here, 1.5 TB (24×64GB) is the priced-but-unselected premium. All three fill every DIMM slot = full 12-channel bandwidth. Pick one in the RAM panel below — it re-prices this line, the roll-up and the returns model live.
Shared infrastructure — one-time · lives in Rack 1, serves both racks
Total $282,000 (street) — includes the HA storage pair, broken-out 400G optics and the production edge-firewall HA pair (selected in the Firewall section — the single biggest swing in this total). The DGX Sparks are costed separately (returns model).
Component / what & why
Part #
Qty
Est.
Lead
Where
RAM — selectable per IT feedback · and how many tenants each config runs
IT review: 768 GB sits exactly at the GPU-VRAM cut-off (8×96 GB = 768 GB), leaving nothing for the OS, KV-cache spill, and page cache — they recommend bumping to 1 TB, with 1.5 TB "even better if money allows… more memory won't hurt the system in the future, only pockets." All three tiers below fill all 24 DIMM slots = full 12-channel bandwidth. Pick one — it re-prices the BOM, headroom and the returns model live. 1.5 TB is priced but not selected.
Why 1 TB is the default now: with 768 GB the host has 0 GB of headroom over the 768 GB of aggregate VRAM — every byte the OS, container runtime, model loader, and KV-cache overflow need comes out of what tenants can stage. 24×48 GB = 1,152 GB keeps full bandwidth and adds ~384 GB of working headroom (≈1.5× VRAM), the cheapest balanced way over the 1 TB line. 1.5 TB (24×64 GB) doubles host RAM to ~2× VRAM — the right call only if you expect large-context or training-style jobs that stage huge datasets in host memory. Note 16×64 GB hits exactly 1 TB but runs only 8 of 12 channels (≈⅓ bandwidth lost) — it is the "1 TB" number without the bandwidth, so it is shown below as avoid.
Config (per server)
Capacity
Channels / bandwidth
Cost / server
Pros & cons
How many tenants can one 1,152 GB / 8-GPU server run?
Typical scenarios:
Bottom line: any of these configs comfortably hosts 8 whole-GPU tenants, or up to ~32 fractional tenants with MIG, or 1 large multi-GPU tenant running a 400B-class model across the box. The difference between them is host-RAM headroom over VRAM — which is exactly the margin IT flagged as missing at 768 GB.
🛡Edge firewall — options, why it's needed & how many servers each fronts
This build puts GPUs on the public internet for paying tenants, reached over a 20 Gb circuit (upgradeable to 100 Gb). The box guarding that edge has to route traffic, wall every tenant off from every other tenant (VDOMs), terminate VPNs and absorb attacks on the public GPU endpoints. The plan only names a placeholder FG-100F — fine for testing, a dead end the moment a paying tenant arrives. Pick the production box below — it flows into the BOM, headroom and returns model. Per IT: the FortiGate 900G in an HA pair is the right-size pick for speed + threat protection "since R&D can turn into 'we're live on this' at any moment."
How to read "servers fronted": inference is light on the edge, so the binding limits are tenant count (VDOMs) and inspection throughput — not server count.
Model traffic streams small tokens and dataset pulls stay on the internal 400 G RoCE fabric, so only tenant-facing internet traffic crosses this box. At a planning figure of ~1.5 Gbps sustained edge egress per 8-GPU server, every production-class box below comfortably fronts the entire container (~24 servers / 192 GPUs) under the planned pass-through posture (TLS terminated at the app layer, not deep-inspected). The inspected figure (if you ever TLS-deep-inspect tenant traffic) is the conservative floor. The real differentiator is VDOMs (how many isolated tenants) and whether the box is 100 G-native for the circuit upgrade.
Model
Strategy
Capex · HA pair (3-yr UTP, street)
Inspected (threat-prot)
Servers fronted pass-through · inspected
Tenants (VDOM)
100 G native
Avg draw
Pick
What changed vs the placeholder: the plan's FG-100F line was a hardware-only ~$4k stand-in. A real production firewall is an HA pair with 3 years of Unified Threat Protection, so the BOM line jumps to the figure for the box you pick (≈$135k for the 900G pair, ≈$248k for the 1000F pair). That capex is a site-wide fixed cost — at 3 servers it is a heavy ~$45k/server burden, but across the full 24-server container it amortises to ~$5–10k/server, which is exactly why the POC IRR looks thin and the full-build IRR (below) is strong. Not in capex: the recurring DDoS-scrubbing + diverse second carrier (~$198–252k over 3 yr) and the ~$4k/mo circuit — those are opex, common to every option.
03Third-party picks (scored)
For each non-Supermicro line: value → premium, with a clickable part # / buy link, specs, street price, and a fit score /10. Voltage: container is 400/230V wye → PDUs 415V WYE, UPS/ATS 230V variants (the listed US GXT5 SKU is 208V — spec the 230V output model).
04Returns — per-asset IRR, scaling & capital recovery
This models one asset class at a time (RTX servers or DGX Sparks), so bundling DGX no longer drags down the RTX number. Set quantity, choose whether the shared infra cost is included, flip 3-yr / 5-yr and salvage on/off, and tune the rate, fee, power and annual rate-decay. Hover any metric for its formula (now in a floating box that won't overlap).
Asset
Quantity
3
Include shared infra in capex?
Horizon & salvage
Rate & utilization
Sky Forge on-demand $1.49
share of hours sold
0% direct · ~20% via Shadeform
Costs & decay
range $0.07–$0.12
added to energy
rental falls each year
Pricing scenarios — one click sets rate · fee · utilization · decay · power. Anchored to the live Sky Forge rate card ($1.49 on-demand · $1.19 reserved).
Capital-recovery schedule — how the APY pays down your principal (yield earned at the IRR each year; balance left after each year's cash)
Year
Balance start
Return earned @ IRR
Cash received
Principal recovered
Balance left
Scaling — what happens as you add RTX servers on the same fixed shared infra (infra cost amortizes → IRR climbs)
RTX servers
Total invested
Loaded $/server
IRR
Payback
3-yr CoC
Pricing scenarios — premium vs base vs heavily-discounted IRR at 3 server(s), current horizon & infra setting (click a row's preset above to load it into the controls)
Scenario
Net $/GPU-hr
Util
Decay
Power
IRR
Payback
Net Y1
3-yr CoC
How these numbers are calculated
Net rate = list rate × (1 − marketplace fee). At $1.75 with a 20% fee you actually collect $1.40/GPU/hr. Revenue (year t) = net rate × 8,760 h × utilization × GPUs × (1 − decay)^(t−1) — so the rental rate steps down each year by the decay %. Energy = kW × 70% load × 1.11 PUE × 8,760 × (power $/kWh + remote-mgmt $/kWh). Net cash = revenue − energy − opex.
IRR solves Σ CFₜ ÷ (1+r)ᵗ = 0 over [−capex, net₁, … , net_N (+salvage)] by bisection. Salvage = 35% of capex in the final year (toggle off to exclude). The capital-recovery schedule treats your IRR as a yield: each year the outstanding balance earns balance × IRR, the cash received first pays that yield and the rest reduces principal — so the balance reaches ~zero in the final year (that's what IRR means). Including shared infra adds its full cost to capex and splits it across your servers (loaded $/server = server cost + infra ÷ N); leaving it out shows the marginal economics of adding a box to infra you already own.
🏗Full build-out — shared infra at scale, the price & the IRR
The Returns scaling table holds shared infra fixed (marginal economics). This view does the opposite: it grows the shared infra with the fleet — adding switches, an HA storage pair, racks, PDUs/UPS and optics as servers go in — so the capex and IRR are the real provisioning numbers. Firewall is a single site-wide HA pair across the whole container. All rows use the base on-demand $1.49 case; flip the scenario presets above to see premium/discounted economics on the same fleet.
Servers
GPUs
Container power
Racks
Shared-infra config (scales)
All-in capex
Loaded $/srv
IRR base · 3-yr
Payback
container at full load — 8 racks, ~24 servers, power vs the 160 kW ceiling
Rack 1 also carries the shared infra (switch, storage pair, head, firewall pair), so it holds 2 GPU servers; the remaining racks hold 3 each. Green bars = GPU servers, cyan = infra.
05What else to add
Most of what used to live here is now in the build. What remains is one Phase-II capital item and two software/strategy upside levers.
Physical security is excluded — you've got that covered. The items below are genuinely optional / deferrable.
06Full-system lead time & cost roll-up
End-to-end timeline if you place POs today, then the combined cost under the active pricing mode. GPUs are not the long pole — DDR5 RAM and the NVIDIA networking are. Place those POs first.