Help me choose

The small startup

You want to run ONE big open model behind your product, as cheaply as it can honestly be run.

Our recommendation: run DeepSeek V4 Flash on 3× H200 NVL

DeepSeek V4 Flash is the biggest bang per dollar on the open leaderboard: frontier-class answers with only 13 billion parameters active per token, which keeps it fast on modest hardware. Three H200 NVL cards bridged by NVLink give you 423 GB — comfortably over its 340.8 GB requirement — for less than a fifth of the price of a DGX.

The memory math behind it

DeepSeek V4 Flash has 284 billion parameters × 1 GB = 284 GB, + 20% working room = 340.8 GB needed — this build carries 423 GB.

The build

3× NVIDIA H200 NVL

141 GB HBM3e each · $32,000 each

The smallest card that takes big models seriously.

$96,000total price
423 GBcombined GPU memory
1,800 Wcombined power draw

1,800 W — as much as about 1.5 average homes drawing power around the clock, and a day of running it uses about 48% of a 90 kWh electric-car battery.

Several cards, one machine — how does that work?

A cluster is several machines (or cards) wired together tightly enough to work on one job as if they were one. These cards are joined by NVLink bridges at 900 GB/s, so the model sees their memory as one pool. That direct wiring is exactly what desktop cards lack — which is why we sell this and not a pile of gaming cards.

I want this build — $96,000

No payment, no account. We save your request, give you a request number, and get back to you.

Want to compare? See all options for this model in the model advisor, or go back to both paths.