gpuos

Use cases · 2 min read · updated Oct 6, 2026

Private RAG on your own GPU: embeddings and chat, self-hosted

Build retrieval-augmented generation that never sends documents to a third party: BGE-M3 embeddings and Qwen3 32B answers on your GPU, called through LangChain.

Why keep RAG private

RAG sends your most sensitive material to the model: contracts, tickets, medical notes, source code. With a hosted API, every chunk you retrieve leaves the company. Running both the embedding model and the chat model on your own GPU keeps the whole loop inside your infrastructure.

Pick the two models

Both fit together on a 24 GB card, so one node serves the whole pipeline.

The pipeline in 20 lines

Install
pip install langchain-openai langchain-core
rag.py
import os
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_openai import ChatOpenAI, OpenAIEmbeddings

BASE, KEY = "https://gpuos.si/v1", os.environ["GPUOS_API_KEY"]

embeddings = OpenAIEmbeddings(model="bge-m3", base_url=BASE, api_key=KEY, check_embedding_ctx_length=False)
llm = ChatOpenAI(model="qwen3-32b", base_url=BASE, api_key=KEY, temperature=0)

docs = [
    "Refunds are accepted within 30 days with the original receipt.",
    "Support is open Monday to Friday, 9:00 to 18:00 CET.",
]
store = InMemoryVectorStore.from_texts(docs, embeddings)

question = "Can I get a refund after three weeks?"
context = "\n\n".join(d.page_content for d in store.similarity_search(question, k=2))
answer = llm.invoke(f"Answer only from this context.\n\n{context}\n\nQuestion: {question}")
print(answer.content)

Swap InMemoryVectorStore for pgvector, Qdrant or any store LangChain supports; the gpuos parts stay the same.

Make it production-ready

  • Use separate API keys for the indexing job and the chat app, with a token quota on the indexing key.
  • Chunk documents to 500 to 1,000 tokens with some overlap; BGE-M3 accepts up to 8K but smaller chunks retrieve more precisely.
  • Watch Usage in the dashboard to see embedding and chat volumes separately.

Questions

Can I use LlamaIndex instead of LangChain?
Yes. Point OpenAILike at the gpuos endpoint for the LLM and use an OpenAI-compatible embedding class for bge-m3.
Does gpuos store my documents?
No. Requests transit through the gpuos gateway to your node; gpuos keeps token counts and latency for metering, not the text.

Related

Run it on your own GPU

Free for one GPU. Connect a machine in one command and call your models through one OpenAI-compatible API.