Layer Zero sells the four rungs of the hardware ladder that enterprise AI models actually run on — from the desktop card that shows you the honest floor, to the full rack that powers a neighborhood.
NVIDIA GeForce RTX 5090
The entry point, and mostly a reality check. At 32 GB it's the cheapest way to get GPU memory, but no single card holds a big model, and desktop cards can't pool. It's here so you can see the honest floor: a pile of these looks cheap but can't cluster to run the models in this store. Useful for understanding why you need datacenter hardware — rarely the thing you actually buy.
No recommended build uses these cards. 11 × 32 GB = 352 GB on paper — but eleven separate islands that can't pool. See the full math on the Flash model page.
Get a quote →NVIDIA H200 NVL
The real building block. 141 GB and, crucially, NVLink — so several pool into one larger machine. The à-la-carte card: you buy the number you need and cluster them yourself (3 for Flash, 7 for GLM). The honest cheapest way to run a big model, if you have the expertise to build and run the cluster. The startup's build.
| Model | Cards | Total memory | Total price | Power |
|---|---|---|---|---|
| DeepSeek V4 Flash | 3 | 423 GB | $107,994 VERIFIED | 1,800 W → 1.5 homes |
| GLM 5.2 | 7 | 987 GB | $251,986 VERIFIED | 4,200 W → 3.5 homes |
| DeepSeek V4 Pro | 14 | 1,974 GB | $503,972 VERIFIED | 8,400 W → 7 homes |
NVIDIA DGX H200
A pre-built machine using eight H200 SXM modules — a server-only version of the chip you can't buy as loose cards — wired and clustered at the factory into one 1,128 GB unit. You reach for this when you'd rather buy one integrated, supported machine than assemble a stack of cards. More than the equivalent loose cards, but it arrives working. The mid-size company's build.
| Model | Units | Total memory | Price | Power |
|---|---|---|---|---|
| DeepSeek V4 Flash | 1× DGX H200 | 1,128 GB | ~$400–500K VERIFIED EXAMPLE | 10.2 kW → 8.5 homes · 2.7 EV batteries/day |
| GLM 5.2 | 1× DGX H200 | 1,128 GB | ~$400–500K VERIFIED EXAMPLE | 10.2 kW → 8.5 homes · 2.7 EV batteries/day |
| DeepSeek V4 Pro | 2× DGX H200 | 2,256 GB | ~$800K–$1M VERIFIED EXAMPLE | 20.4 kW → 17 homes · 5.4 EV batteries/day |
NVIDIA GB300 NVL72
The top — 72 GPUs acting as a single ~20 TB machine, drawing as much power as ~100 homes. For running the largest models at scale, or many models at once. For a single model it's overkill by design. The "when you outgrow servers entirely" ceiling — if you have the facility and budget to match.
| Model | Units | % of rack used | Price | Power |
|---|---|---|---|---|
| Any model | 1× GB300 NVL72 | Flash ~2% · GLM ~4% · Pro ~10% | ~$3.7–6.5M ESTIMATE No official NVIDIA list price. | 120 kW → 100 homes · 32 EV batteries/day |
For a single model, this rack is overkill. It makes sense when running multiple large models simultaneously or at the highest inference volumes.