Point the client at gpuos
The gpuos gateway speaks the OpenAI API, so the official openai package works as is. Set base_url to your gpuos endpoint and use a gpuos key.
pip install openaiimport os
from openai import OpenAI
client = OpenAI(base_url="https://gpuos.si/v1", api_key=os.environ["GPUOS_API_KEY"])
reply = client.chat.completions.create(
model="qwen3-32b",
messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
)
print(reply.choices[0].message.content)Create the key in API keys in your gpuos dashboard, and deploy the model first in Models. GET /v1/models lists what your workspace can call.
Streaming and embeddings
stream = client.chat.completions.create(
model="qwen3-32b",
messages=[{"role": "user", "content": "Explain VRAM in two sentences."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
vectors = client.embeddings.create(model="bge-m3", input=["first text", "second text"])Errors come back in the OpenAI format, so openai.RateLimitError, openai.AuthenticationError and openai.PermissionDeniedError behave as usual. A quota or rate limit set on the key returns a 429.
Official documentation: github.com/openai/openai-python