gpuos

Hardware · 2 min read · updated Oct 6, 2026

Run Qwen3 32B on a Hetzner GPU server (GEX45)

Set up a Hetzner GEX45 with an RTX PRO 4000 Blackwell 24 GB, install the NVIDIA driver and gpuos, and serve Qwen3 32B through an OpenAI-compatible API.

The server

Hetzner's GEX45 is a dedicated server with an NVIDIA RTX PRO 4000 Blackwell SFF, 24 GB of ECC GDDR7, an Intel i5-13500, 64 GB of RAM and two 512 GB NVMe drives, hosted in the EU. Check Hetzner’s current offer for price and setup fees.

24 GB is the sweet spot for 32B models at 4-bit, which is where open models become good enough for most internal work.

Prepare the machine

  1. Order the GEX45 with Ubuntu 24.04. New Hetzner accounts can face a manual check for GPU servers, so order early.
  2. Install the NVIDIA driver, then reboot.
  3. Check that the GPU is visible with nvidia-smi.
On the server
sudo apt update && sudo apt install -y ubuntu-drivers-common
sudo ubuntu-drivers install
sudo reboot
# after the reboot
nvidia-smi

Install gpuos and deploy Qwen3 32B

  1. In your gpuos workspace, open Nodes → Add node and copy the install command.
  2. Run it on the server. Within seconds the node shows up with its RTX PRO 4000 and 24 GB of VRAM.
  3. Open Models, choose Qwen3 32B: the gauge shows about 22 GB of 24 GB. Confirm.
  4. Create an API key and call the model.
Python
from openai import OpenAI

client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
reply = client.chat.completions.create(
    model="qwen3-32b",
    messages=[{"role": "user", "content": "What can you do on a 24 GB GPU?"}],
)
print(reply.choices[0].message.content)

Tips for this box

  • Disk is the tightest resource: a 32B model takes about 20 GB on disk. Keep the catalog small and remove models you no longer use from Models.
  • For coding assistants, add gpt-oss 20B next to Qwen3 32B; Ollama swaps them when both do not fit at once.
  • When you outgrow 24 GB, the GEX131 (96 GB) runs 70B models; add it as a second node in the same workspace.

Verify GPU loading before serving a team

Check the current GEX45 specification before ordering; region availability, price and setup fees can change. This guide uses the 24 GB RTX PRO 4000 Blackwell SFF configuration.

Check the GPU and running model
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
ollama --version
ollama show qwen3:32b
ollama ps

After a request loads Qwen3, ollama ps shows the running model and processor allocation. Confirm GPU offloading, the quantization and the requested context. The Qwen3-32B page includes the local setup command.

  • If nvidia-smi cannot see the card, fix the driver and reboot before debugging the gateway.
  • If inference uses CPU memory, inspect the runtime’s GPU support and free VRAM. A running service alone does not prove GPU inference.
  • Qwen3-32B leaves little room on 24 GB. Start with a short context and one request; test longer prompts and concurrent requests separately.

The fit estimate is available in the VRAM calculator. No tokens-per-second result is claimed here: throughput depends on the driver, Ollama version, context and workload.

Questions

Does the GEX45 need a public port open for gpuos?
No. The agent connects out over HTTPS, so you can keep Hetzner's firewall closed to inbound traffic except SSH.
Is the GEX45 enough for a team?
For a small team, yes: a 32B model serves several users at interactive speed. Use per-key rate limits to share it fairly.

Related

Run it on your own GPU

Free for one GPU. Connect a machine in one command and call your models through one OpenAI-compatible API.