gpuos

Integration · Proxy

LiteLLM with your own GPUs

Add your own GPUs to a LiteLLM proxy next to cloud providers, with gpuos as an OpenAI-compatible backend.

Base URL
https://gpuos.si/v1
API key
gpuos_key_… from the dashboard
Model
a catalog id, e.g. qwen3-32b

Add gpuos models to LiteLLM

config.yaml
model_list:
  - model_name: qwen3-32b
    litellm_params:
      model: openai/qwen3-32b
      api_base: https://gpuos.si/v1
      api_key: os.environ/GPUOS_API_KEY
  - model_name: bge-m3
    litellm_params:
      model: openai/bge-m3
      api_base: https://gpuos.si/v1
      api_key: os.environ/GPUOS_API_KEY

LiteLLM then routes these model names to your GPUs and everything else to the providers you already use.

Official documentation: docs.litellm.ai

Questions

Why put gpuos behind LiteLLM?
If your apps already use LiteLLM, this adds your own GPUs as one more backend without changing client code. gpuos handles the GPU side: nodes, model deployment and per-key metering.
Can LiteLLM fall back to a cloud model if my node is offline?
Yes. gpuos returns a 503 when no node serving the model is online, which LiteLLM fallbacks can catch.

Related

Use LiteLLM with models on your GPUs

Free for one GPU. Connect a machine, deploy a model, then paste the base URL and your key.