The price on the card is the price on the invoice.
No setup fee. No egress billing. No support tier to upgrade into. One number a month for the whole machine, cancel with 30 days' notice.
Ryzen AI Max+ 395
128 GB shared between CPU and GPU — a 70B at 8-bit, or a 120B-class model at 4-bit, in memory that would cost 5× this in discrete cards.
One 128 GB pool shared by CPU and GPU — that's the whole point of this box. The catch is bandwidth: 256 GB/s means tokens generate slower than on a discrete card, so it suits large models, batch work and dev rather than high-throughput serving. Runs on ROCm, not CUDA.
2× AMD Radeon AI PRO R9700
A 70B at 4-bit across both cards, or a 32B at fp8 with plenty of context headroom.
Runs on ROCm, not CUDA. PyTorch, vLLM, llama.cpp and Ollama all work — but check your stack before you commit, and ask us if you're unsure.
RTX 5090
A 32B at fp8 with real context, or SDXL / Flux / ComfyUI without queueing.
2× RTX 5090
A 70B at 4-bit across both cards, or two independent 32B endpoints.
RTX PRO 6000 Blackwell
A 70B at fp8 with long context on ONE card — no tensor-parallel setup to debug.
2× RTX PRO 6000 Blackwell
A 120B-class open model at fp8, or a full fine-tune of a 35B.
Every plan: full root on bare metal · dedicated card, nothing shared · 1 dedicated IPv4 · Ubuntu, Windows or your own ISO
Extras, priced up front.
Where we win, and where we don't.
Including the row where we lose. If per-second billing is what you need, we'll tell you to go elsewhere rather than sell you the wrong thing.
| GPURack | Hyperscaler | Typical GPU host | |
|---|---|---|---|
| Dedicated card, nothing shared | |||
| Full root on bare metal | |||
| Flat monthly rate | |||
| No egress / bandwidth billing | |||
| Two independent fiber carriers | |||
| Solar + battery behind the load | |||
| Talk to the person who built it | |||
| Per-second billing We're honest: if your workload is genuinely bursty, spot pricing beats us. |
means it depends on the provider and the plan.
The things people ask first.
Why is this cheaper than AWS or Google Cloud?
What actually happens during a power outage?
Is the GPU shared with anyone else?
Which card do I need for my model?
Do I get root access? Can I run Docker?
You rent AMD boxes. Does my stack actually run on them?
How fast is deployment?
Do you offer hourly billing?
What's the contract and how do I pay?
Who do I talk to when something breaks?
Tell us what you're running.
Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.
No setup fee · Month to month · sales@gpurack.net