Configure the LLM
OpenAILike is LlamaIndex's client for OpenAI-compatible servers. Set is_chat_model=True so it calls /chat/completions.
pip install llama-index-llms-openai-likeimport os
from llama_index.core import Settings
from llama_index.llms.openai_like import OpenAILike
Settings.llm = OpenAILike(
model="qwen3-32b",
api_base="https://gpuos.si/v1",
api_key=os.environ["GPUOS_API_KEY"],
is_chat_model=True,
context_window=32768,
)Create the key in API keys in your gpuos dashboard, and deploy the model first in Models. GET /v1/models lists what your workspace can call.
Official documentation: www.llamaindex.ai