GRGPURackAI infrastructureGet a box
Pricing

The price on the card is the price on the invoice.

No setup fee. No egress billing. No support tier to upgrade into. One number a month for the whole machine, cancel with 30 days' notice.

Starter
ROCm

Ryzen AI Max+ 395

128 GB unified LPDDR5X

128 GB shared between CPU and GPU — a 70B at 8-bit, or a 120B-class model at 4-bit, in memory that would cost 5× this in discrete cards.

Stream processors2,560 (40 CU)
FP3229.7 TFLOPS peak
Mem bandwidth256 GB/s
CPU16C / 32T Zen 5
RAM128 GB LPDDR5X (shared CPU+GPU)
Storage2 TB NVMe
Network1 Gbps unmetered

One 128 GB pool shared by CPU and GPU — that's the whole point of this box. The catch is bandwidth: 256 GB/s means tokens generate slower than on a discrete card, so it suits large models, batch work and dev rather than high-throughput serving. Runs on ROCm, not CUDA.

The most addressable memory per dollar we rent
$249/month
Reserve this box
2 available now
Standard
ROCm

AMD Radeon AI PRO R9700

64 GB total GDDR6

A 70B at 4-bit across both cards, or a 32B at fp8 with plenty of context headroom.

Stream processors4,096 per card
FP3247.8 TFLOPS per card
Mem bandwidth640 GB/s per card
CPU24 vCPU
RAM128 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered

Runs on ROCm, not CUDA. PyTorch, vLLM, llama.cpp and Ollama all work — but check your stack before you commit, and ask us if you're unsure.

The cheapest route into 64 GB of VRAM
$399/month
Reserve this box
1 available now
Most popular
Standard
CUDA

RTX 5090

32 GB GDDR7

A 32B at fp8 with real context, or SDXL / Flux / ComfyUI without queueing.

CUDA cores21,760
FP32105 TFLOPS
Mem bandwidth1.79 TB/s
CPU24 vCPU
RAM128 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered
Production inference, image and video generation
$449/month
Reserve this box
3 available now
Multi-GPU
CUDA

RTX 5090

64 GB total GDDR7

A 70B at 4-bit across both cards, or two independent 32B endpoints.

CUDA cores21,760 per card
FP32105 TFLOPS per card
Mem bandwidth1.79 TB/s per card
CPU32 vCPU
RAM192 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered
Tensor-parallel serving, parallel render pipelines
$849/month
Reserve this box
1 available now
Pro
CUDA

RTX PRO 6000 Blackwell

96 GB GDDR7

A 70B at fp8 with long context on ONE card — no tensor-parallel setup to debug.

CUDA cores24,064
FP32126 TFLOPS
Mem bandwidth1.8 TB/s
CPU32 vCPU
RAM256 GB DDR5
Storage4 TB NVMe
Network1 Gbps unmetered
Large-model serving on a single card, LoRA fine-tuning
$949/month
Reserve this box
1 available now
Multi-GPU
CUDA

RTX PRO 6000 Blackwell

192 GB total GDDR7

A 120B-class open model at fp8, or a full fine-tune of a 35B.

CUDA cores24,064 per card
FP32126 TFLOPS per card
Mem bandwidth1.8 TB/s per card
CPU48 vCPU
RAM512 GB DDR5
Storage8 TB NVMe
Network1 Gbps unmetered
Frontier open-weight models, multi-tenant inference
$1,799/month
Reserve this box

Every plan: full root on bare metal · dedicated card, nothing shared · 1 dedicated IPv4 · Ubuntu, Windows or your own ISO

Add-ons

Extras, priced up front.

Additional 1 TB NVMe$25/mo
Additional IPv4$3/mo
10 Gbps uplink$45/mo
Nightly off-site backupfrom $30/mo
Managed AI deploymentfrom $400/mo
Windows Server licence$25/mo
How we compare

Where we win, and where we don't.

Including the row where we lose. If per-second billing is what you need, we'll tell you to go elsewhere rather than sell you the wrong thing.

 GPURackHyperscalerTypical GPU host
Dedicated card, nothing shared
Full root on bare metal
Flat monthly rate
No egress / bandwidth billing
Two independent fiber carriers
Solar + battery behind the load
Talk to the person who built it
Per-second billing
We're honest: if your workload is genuinely bursty, spot pricing beats us.

means it depends on the provider and the plan.

Questions

The things people ask first.

Why is this cheaper than AWS or Google Cloud?
Because you're renting the actual machine, not a slice of one with a hyperscaler's margin stacked on top. There's no egress billing, no per-hour meter, no support tier to upgrade into. One flat monthly number for the whole box, and the GPU is yours alone for the month.
What actually happens during a power outage?
Nothing you'd notice. The facility runs behind a battery bank with a solar array feeding it, so a utility cut is a transfer, not an interruption — the batteries carry the load and the panels recharge them. Network is the same story: two independent fiber carriers, not two drops from one provider, so a single carrier's outage doesn't take the rack with it. This is the reason the business exists.
Is the GPU shared with anyone else?
No. Every plan is a dedicated card in a dedicated box. You get the full VRAM, the full clock, and root. Nothing is virtualised, oversold, or time-sliced, so throughput doesn't move around based on who else is on the machine.
Which card do I need for my model?
Rough rule for inference: VRAM in GB should be about double the parameter count in billions at fp16, or roughly equal to it at fp8. A 32B fits an RTX 5090 at fp8; a 70B wants the 96 GB RTX PRO 6000 to stay on one card. Tell us the model and the context length you're targeting and we'll tell you honestly which box to take — including when the cheaper one is enough.
Do I get root access? Can I run Docker?
Full root on bare metal, and yes to Docker. Ubuntu 24.04 by default with the driver stack already in place — CUDA on the NVIDIA boxes, ROCm on the AMD ones — plus the container toolkit. vLLM, Ollama, llama.cpp, ComfyUI, PyTorch and TensorFlow all run as they do on any other machine you own. Bring your own ISO if you'd rather.
You rent AMD boxes. Does my stack actually run on them?
Usually, but check first — this is the one thing worth five minutes before you order. PyTorch, vLLM, llama.cpp and Ollama all have working ROCm support, and for straightforward LLM inference the AMD boxes are the cheapest VRAM we rent. Where it gets thin is custom CUDA kernels, some quantisation libraries, and anything depending on a niche NVIDIA-only package. Tell us what you're running and we'll give you a straight answer rather than a maybe — and if the answer is that you need CUDA, we'll point you at the NVIDIA plans instead.
How fast is deployment?
In-stock configurations are typically handed over the same day — you get an IP, root credentials, and a working driver stack. Custom builds depend on parts and we'll give you a real date rather than a hopeful one.
Do you offer hourly billing?
Not today. Hourly billing means keeping cards idle between customers, and the cost of that idle time gets priced back into every hour you do use. Flat monthly is how the rate stays where it is. If your workload is genuinely bursty, a hyperscaler's spot market will beat us and we'll say so.
What's the contract and how do I pay?
Month to month. No setup fee, no minimum term, cancel with 30 days' notice. Invoiced monthly by card or ACH.
Who do I talk to when something breaks?
The person who built the rack. There's no ticket tier and no offshore first line — you get a direct line to someone who can actually get hands on the hardware. That's a genuine advantage of being small, and it stops being true if we ever get big enough to need a call centre.

Tell us what you're running.

Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.

No setup fee · Month to month · sales@gpurack.net