gpuos

Integration · Framework

LlamaIndex with your own GPUs

Use gpuos models in LlamaIndex with the OpenAILike LLM and an OpenAI-compatible embedding model.

Base URL
https://gpuos.si/v1
API key
gpuos_key_… from the dashboard
Model
a catalog id, e.g. qwen3-32b

Configure the LLM

OpenAILike is LlamaIndex's client for OpenAI-compatible servers. Set is_chat_model=True so it calls /chat/completions.

Install
pip install llama-index-llms-openai-like
Python
import os
from llama_index.core import Settings
from llama_index.llms.openai_like import OpenAILike

Settings.llm = OpenAILike(
    model="qwen3-32b",
    api_base="https://gpuos.si/v1",
    api_key=os.environ["GPUOS_API_KEY"],
    is_chat_model=True,
    context_window=32768,
)

Create the key in API keys in your gpuos dashboard, and deploy the model first in Models. GET /v1/models lists what your workspace can call.

Official documentation: www.llamaindex.ai

Questions

Why OpenAILike and not the OpenAI class?
The OpenAI class validates model names against OpenAI's list. OpenAILike accepts any model id, such as qwen3-32b.
What context window should I set?
Use the model's context from the gpuos catalog, for example 32768 for Qwen3 32B or 131072 for gpt-oss 20B.

Related

Use LlamaIndex with models on your GPUs

Free for one GPU. Connect a machine, deploy a model, then paste the base URL and your key.