gpuos

Integration · Framework

LangChain with your own GPUs

ChatOpenAI and OpenAIEmbeddings work with gpuos models for chains, agents and RAG.

Base URL
https://gpuos.si/v1
API key
gpuos_key_… from the dashboard
Model
a catalog id, e.g. qwen3-32b

Chat models

Install
pip install langchain-openai
Python
import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="qwen3-32b",
    base_url="https://gpuos.si/v1",
    api_key=os.environ["GPUOS_API_KEY"],
    temperature=0.2,
)
print(llm.invoke("Give me three names for an internal AI assistant.").content)

Embeddings

Use bge-m3 for embeddings. Turn off check_embedding_ctx_length, which assumes OpenAI's tokenizer and would split inputs with the wrong token counts.

Python
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(
    model="bge-m3",
    base_url="https://gpuos.si/v1",
    api_key=os.environ["GPUOS_API_KEY"],
    check_embedding_ctx_length=False,
)

For a full example with a vector store, see Private RAG on your own GPU.

Official documentation: python.langchain.com

Questions

Does tool calling work with LangChain agents?
It depends on the model. Models trained for function calling, such as GLM-4 32B and Qwen3, handle bind_tools well; small models are less reliable.
Is langchain-openai the right package, not a community one?
Yes. langchain-openai accepts any OpenAI-compatible base_url, which is exactly what gpuos exposes.

Related

Use LangChain with models on your GPUs

Free for one GPU. Connect a machine, deploy a model, then paste the base URL and your key.