The server
Hetzner's GEX45 is a dedicated server with an NVIDIA RTX PRO 4000 Blackwell SFF, 24 GB of ECC GDDR7, an Intel i5-13500, 64 GB of RAM and two 512 GB NVMe drives, hosted in the EU. Check Hetzner’s current offer for price and setup fees.
24 GB is the sweet spot for 32B models at 4-bit, which is where open models become good enough for most internal work.
Prepare the machine
- Order the GEX45 with Ubuntu 24.04. New Hetzner accounts can face a manual check for GPU servers, so order early.
- Install the NVIDIA driver, then reboot.
- Check that the GPU is visible with
nvidia-smi.
sudo apt update && sudo apt install -y ubuntu-drivers-common
sudo ubuntu-drivers install
sudo reboot
# after the reboot
nvidia-smiInstall gpuos and deploy Qwen3 32B
- In your gpuos workspace, open Nodes → Add node and copy the install command.
- Run it on the server. Within seconds the node shows up with its RTX PRO 4000 and 24 GB of VRAM.
- Open Models, choose Qwen3 32B: the gauge shows about 22 GB of 24 GB. Confirm.
- Create an API key and call the model.
from openai import OpenAI
client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
reply = client.chat.completions.create(
model="qwen3-32b",
messages=[{"role": "user", "content": "What can you do on a 24 GB GPU?"}],
)
print(reply.choices[0].message.content)Tips for this box
- Disk is the tightest resource: a 32B model takes about 20 GB on disk. Keep the catalog small and remove models you no longer use from Models.
- For coding assistants, add gpt-oss 20B next to Qwen3 32B; Ollama swaps them when both do not fit at once.
- When you outgrow 24 GB, the GEX131 (96 GB) runs 70B models; add it as a second node in the same workspace.
Verify GPU loading before serving a team
Check the current GEX45 specification before ordering; region availability, price and setup fees can change. This guide uses the 24 GB RTX PRO 4000 Blackwell SFF configuration.
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
ollama --version
ollama show qwen3:32b
ollama psAfter a request loads Qwen3, ollama ps shows the running model and processor allocation. Confirm GPU offloading, the quantization and the requested context. The Qwen3-32B page includes the local setup command.
- If
nvidia-smicannot see the card, fix the driver and reboot before debugging the gateway. - If inference uses CPU memory, inspect the runtime’s GPU support and free VRAM. A running service alone does not prove GPU inference.
- Qwen3-32B leaves little room on 24 GB. Start with a short context and one request; test longer prompts and concurrent requests separately.
The fit estimate is available in the VRAM calculator. No tokens-per-second result is claimed here: throughput depends on the driver, Ollama version, context and workload.