Why keep RAG private
RAG sends your most sensitive material to the model: contracts, tickets, medical notes, source code. With a hosted API, every chunk you retrieve leaves the company. Running both the embedding model and the chat model on your own GPU keeps the whole loop inside your infrastructure.
Pick the two models
The pipeline in 20 lines
pip install langchain-openai langchain-coreimport os
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
BASE, KEY = "https://gpuos.si/v1", os.environ["GPUOS_API_KEY"]
embeddings = OpenAIEmbeddings(model="bge-m3", base_url=BASE, api_key=KEY, check_embedding_ctx_length=False)
llm = ChatOpenAI(model="qwen3-32b", base_url=BASE, api_key=KEY, temperature=0)
docs = [
"Refunds are accepted within 30 days with the original receipt.",
"Support is open Monday to Friday, 9:00 to 18:00 CET.",
]
store = InMemoryVectorStore.from_texts(docs, embeddings)
question = "Can I get a refund after three weeks?"
context = "\n\n".join(d.page_content for d in store.similarity_search(question, k=2))
answer = llm.invoke(f"Answer only from this context.\n\n{context}\n\nQuestion: {question}")
print(answer.content)Swap InMemoryVectorStore for pgvector, Qdrant or any store LangChain supports; the gpuos parts stay the same.
Make it production-ready
- Use separate API keys for the indexing job and the chat app, with a token quota on the indexing key.
- Chunk documents to 500 to 1,000 tokens with some overlap; BGE-M3 accepts up to 8K but smaller chunks retrieve more precisely.
- Watch Usage in the dashboard to see embedding and chat volumes separately.