gpuos

Integration · Chat interface

Connect Open WebUI to a self-hosted OpenAI-compatible API

Give your team a ChatGPT-style interface on top of models running on your GPUs.

Base URL
https://gpuos.si/v1
API key
gpuos_key_… from the dashboard
Model
a catalog id, e.g. qwen3-32b

Add gpuos as an OpenAI connection

  1. Open Admin Panel → Settings → Connections.
  2. Under OpenAI API, add a connection with URL https://gpuos.si/v1 and your gpuos key.
  3. Save. The models deployed in your workspace appear in the model picker.

You can also configure it with environment variables when you start the container:

Docker
docker run -d -p 3000:8080 \
  -e OPENAI_API_BASE_URL=https://gpuos.si/v1 \
  -e OPENAI_API_KEY=$GPUOS_API_KEY \
  -v open-webui:/app/backend/data \
  ghcr.io/open-webui/open-webui:main

Official documentation: openwebui.com

Questions

Why not point Open WebUI at Ollama directly?
You can on a single machine. Going through gpuos adds keys, quotas, usage metering and access to models on several nodes, without exposing Ollama's port.
Which model should be the default?
Qwen3 32B or GLM-4 32B on a 24 GB GPU; Qwen3 8B on smaller cards.

Related

Use Open WebUI with models on your GPUs

Free for one GPU. Connect a machine, deploy a model, then paste the base URL and your key.