DeepSeek V4 Flash
An efficiency-first mixture-of-experts model: 284 billion parameters in total, but only 13 billion do the work on each word it produces. All 284 billion still have to sit in GPU memory — the model chooses which experts to wake per token, so every expert must already be loaded.
284 billion parameters × 1 GB = 284 GB → + 20% working room = 340.8 GB of GPU memory
3× H200 NVL
The minimum honest setup: 423 GB of NVLink-bridged memory against the 340.8 GB the model needs.
1,800 W — as much as about 1.5 average homes drawing power around the clock, and a day of running it uses about 48% of a 90 kWh electric-car battery.
Request this build — $96,000
DGX B300
The room-to-grow pick: the same model with 6.7× headroom for more users, longer contexts, or a bigger model later.
14,000 W — as much power as about 12 average homes, and a day of running it drains about 3.7 full 90 kWh electric-car batteries.