Guides
Everything we learned running open models on our own GPUs: what fits in VRAM, how to expose a model safely, and how to build private SI features on top.
Getting started
Hardware
- How much VRAM do you need to run an LLM?VRAM needed for Qwen3, GLM-4, gpt-oss, DeepSeek R1 distill, Mistral Small, Gemma 3 and more, the rule of thumb behind it, and which GPUs fit each model.2 min read
- Run Qwen3 32B on a Hetzner GPU server (GEX45)Set up a Hetzner GEX45 with an RTX PRO 4000 Blackwell 24 GB, install the NVIDIA driver and gpuos, and serve Qwen3 32B through an OpenAI-compatible API.2 min read
Use cases
Compliance
Comparisons
Try it on your own GPU
Free for one GPU. Install the agent, deploy a model and call it through one OpenAI-compatible API.