Choose an embedding model for retrieval, not chat
An embeddings API turns text into vectors you can compare for semantic retrieval. BGE-M3 and Qwen3 Embedding 0.6B are two catalog candidates for this stage of a RAG pipeline. They do not generate the final answer; use a chat model after retrieving relevant passages.
This tutorial builds a small, runnable dense-retrieval client with Python's standard library. Run it against local Ollama first, then use the same request shape with a deployed gpuOS model. The fixture demonstrates API handling and ranking; it is not a benchmark of either model.
Scroll horizontally to see every column.
| Model | Local Ollama tag | Catalog quantization | Estimated VRAM | Catalog context (tokens) | License |
|---|---|---|---|---|---|
| BGE-M3 | bge-m3 | FP16 | 1.5 GB | 8,192 | MIT |
| Qwen3 Embedding 0.6B | qwen3-embedding:0.6b | Q8_0 | 2 GB | 32,768 | Apache 2.0 |
Memory values are planning estimates for short inputs. Long chunks and large batches can raise peak memory. The context column is an input limit, not a recommended chunk size or a promise that a maximum-size batch fits.
The differences that change your client and index
Scroll horizontally to see every column.
| Decision | BGE-M3 | Qwen3 Embedding 0.6B |
|---|---|---|
| Query formatting | No retrieval instruction required | Prefix queries with a task instruction; leave documents unprefixed |
| Publisher vector dimension | 1024 | Up to 1024; the publisher supports smaller dimensions |
| API recipe here | One dense vector per text | One dense vector per text, default dimensions |
| Switching models | Rebuild document vectors and re-evaluate retrieval | Rebuild document vectors and re-evaluate retrieval |
The BGE-M3 publisher also provides sparse and multi-vector retrieval through its own tooling. This OpenAI-compatible embeddings recipe uses dense vectors only; selecting BGE-M3 does not automatically create a hybrid retriever. BAAI model card.
Qwen's query instruction describes the retrieval task; it is not an instruction for a chat assistant. The script adds it only to the query. It leaves output dimensions at the server default and reads the actual response size rather than assuming a custom-dimension API works in every runtime. Qwen model card.
Start an API on the machine running Ollama
Prerequisites: Python 3, an installed and running Ollama server, enough RAM/VRAM for your selected model, and disk space for its weights. Keep the default local bind for this test. Pull only the candidate you plan to evaluate first.
ollama --version
ollama pull bge-m3
curl --fail-with-body http://127.0.0.1:11434/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model":"bge-m3","input":["Rotate a leaked API key.","VRAM stores GPU model weights."],"encoding_format":"float"}'Expect a JSON data array with one indexed embedding vector per input. Ollama's native /api/embed returns a different shape with an embeddings array. Choose one protocol for your client; the script below uses /v1/embeddings. Ollama OpenAI compatibility.
For Qwen, pull qwen3-embedding:0.6b and pass that exact local tag. A connection refusal means the local server is unavailable; a model-not-found response means the selected tag has not been downloaded or does not match.
Run a complete dense-retrieval example
Save the following as retrieve.py. It embeds three short passages and one question, restores the API's indexed order, checks numeric vectors and computes cosine similarity. It runs with no pip dependencies. The client refuses redirects; configure the final API URL directly so the key stays on the intended endpoint.
import json
import math
import os
import urllib.error
import urllib.request
BASE = os.environ.get("EMBEDDING_BASE_URL", "http://127.0.0.1:11434/v1").rstrip("/")
MODEL = os.environ.get("EMBEDDING_MODEL", "bge-m3")
KEY = os.environ.get("GPUOS_API_KEY", "")
DOCUMENTS = [
("gpu", "The GPU node runs the model. Its VRAM must fit weights and the workload."),
("key", "An API key authorizes access. Revoke a leaked key and issue a replacement."),
("index", "An embedding index stores document vectors for semantic retrieval."),
]
QUESTION = "What should I do if an API key leaks?"
class NoRedirectHandler(urllib.request.HTTPRedirectHandler):
def redirect_request(self, request, fp, code, message, headers, new_url):
return None
# Refuse redirects so an Authorization header cannot reach a different endpoint.
CLIENT = urllib.request.build_opener(NoRedirectHandler())
def query_text(question):
if MODEL.startswith("qwen3-embedding"):
return "Instruct: Retrieve the passage that answers the question.\nQuery:" + question
return question
def embed(texts):
if any(ord(character) <= 32 or ord(character) >= 127 for character in KEY):
raise ValueError("API key must contain printable ASCII characters without spaces")
headers = {"Content-Type": "application/json"}
if KEY:
headers["Authorization"] = "Bearer " + KEY
request = urllib.request.Request(
BASE + "/embeddings",
data=json.dumps({"model": MODEL, "input": texts, "encoding_format": "float"}).encode(),
headers=headers,
method="POST",
)
with CLIENT.open(request, timeout=120) as response:
payload = json.load(response)
if not isinstance(payload, dict):
raise ValueError("Expected an embedding response object")
rows = payload.get("data")
if not isinstance(rows, list) or len(rows) != len(texts):
raise ValueError("Expected one embedding row per input")
by_index = {}
for row in rows:
if not isinstance(row, dict):
raise ValueError("Expected an embedding row object")
index = row.get("index")
vector = row.get("embedding")
if type(index) is not int or not 0 <= index < len(texts) or index in by_index:
raise ValueError("Invalid or duplicate embedding index")
if not isinstance(vector, list) or not vector or any(
type(value) not in (int, float) or not math.isfinite(value) for value in vector
):
raise ValueError("Expected a nonempty finite numeric vector")
by_index[index] = vector
vectors = [by_index[index] for index in range(len(texts))]
if len({len(vector) for vector in vectors}) != 1:
raise ValueError("Embedding dimensions differ within the batch")
return vectors
def cosine(left, right):
if len(left) != len(right):
raise ValueError("Embedding dimensions differ")
denominator = math.sqrt(sum(x * x for x in left) * sum(x * x for x in right))
if denominator == 0:
raise ValueError("Cannot rank a zero vector")
return sum(a * b for a, b in zip(left, right)) / denominator
def main():
vectors = embed([text for _, text in DOCUMENTS] + [query_text(QUESTION)])
ranking = sorted(
((cosine(vector, vectors[-1]), doc_id, text)
for (doc_id, text), vector in zip(DOCUMENTS, vectors[:-1])),
reverse=True,
)
print("model=" + MODEL, "dimensions=" + str(len(vectors[0])))
for score, doc_id, text in ranking:
print(f"{score:.4f} {doc_id}: {text}")
if __name__ == "__main__":
try:
main()
except urllib.error.HTTPError as error:
raise SystemExit(f"Embedding HTTP {error.code}: check model, key and server logs")
except (urllib.error.URLError, TimeoutError, ValueError, KeyError, TypeError) as error:
raise SystemExit(f"Embedding request failed: {error}")
python3 retrieve.py
# Compare the second candidate with the same document fixture.
ollama pull qwen3-embedding:0.6b
EMBEDDING_MODEL=qwen3-embedding:0.6b python3 retrieve.pyThe output prints the model, actual dimension and three ranked passages. For this question, the key passage is the relevant item to check. Scores and ordering depend on the model and runtime; there is no fixed expected score. A valid response with poor ranking is a retrieval-quality issue, not an HTTP success criterion.
This tiny fixture is an API smoke test. For a real index, embed documents during ingestion, store their vectors and metadata, then embed only the user's query during retrieval. Do not re-embed the whole corpus on every request.
Use the same client with a deployed gpuOS model
Connect a GPU node through the quickstart, deploy BGE-M3 or Qwen3 Embedding 0.6B, and create a server-side workspace API key. Both are currently Community catalog models. Set GPUOS_API_KEY through your environment or secret manager before running this command.
# GPUOS_API_KEY must already be set in the environment.
EMBEDDING_BASE_URL=https://gpuos.si/v1 \
EMBEDDING_MODEL=bge-m3 \
python3 retrieve.py
# Deploy Qwen3 Embedding first, then use its catalog ID here.
EMBEDDING_BASE_URL=https://gpuos.si/v1 \
EMBEDDING_MODEL=qwen3-embedding-0.6b \
python3 retrieve.pygpuOS accepts catalog IDs such as qwen3-embedding-0.6b; local Ollama uses tags such as qwen3-embedding:0.6b. The gateway forwards embeddings requests to the connected node's Ollama route. Your input passages and returned vectors pass through the hosted gpuOS gateway, which authorizes requests and meters usage. This is not an entirely local or air-gapped request path.
- 401: check the workspace key and Bearer header; keep keys out of browser bundles.
- 404: check the gpuOS catalog ID rather than the Ollama tag.
- 429: inspect the key's request limit or monthly token quota before retrying.
- 503: confirm that the model is deployed and its node is connected.
- 400 or 502 with an engine error: inspect the upstream input/batch limits and runtime logs; reduce the batch or chunk size before trying again.
Compare retrieval quality, latency and memory
- Build a labeled set from actual queries, including abbreviations, multilingual inputs and queries with no answer in the corpus. Keep the same documents and chunk boundaries for both candidates.
- Measure whether an expected passage appears in the first k results (recall at k). Inspect incorrect top results and near-duplicate chunks instead of comparing raw cosine scores across models.
- Record query latency separately from ingestion throughput. Measure cold model loading and warm requests separately, and vary batch size on the hardware you will use.
- Record peak memory while embedding realistic chunks. If a maximum-size input fails or truncates, shorten the chunks and re-evaluate answer coverage.
- Store the model/tag or digest, quantization, actual vector dimension, query-instruction template and chunking version with the index. Keep this configuration fixed during evaluation.
Embedding scores are ranking values, not calibrated probabilities. Set any relevance threshold using your own positive and negative examples. For RAG, retrieve the passages first and pass their text plus source IDs to a chat model using the private RAG guide.
When testing the native Ollama embedding API, truncate:false lets you reject inputs above its context limit instead of silently shortening them. That is an /api/embed option; do not assume the OpenAI-compatible endpoint accepts the same option. Native embedding request options.
Change models without corrupting the index
The same vector dimension does not make two embedding spaces interchangeable. If you move from BGE-M3 to Qwen3 Embedding, build a new document index, evaluate it with Qwen-formatted queries, then switch both ingestion and retrieval together. Keep the previous index until you can roll back both sides.
Never mix old document vectors with new query vectors or append a new model's vectors into an existing collection. For instruction or chunking changes, keep a separate evaluation/version as well. Plan the backfill time, disk space, access-control metadata and deletion behavior before moving a production corpus.
Use LlamaIndex or LangChain when your application needs document loaders and a persistent vector store. Keep the embedding model and query formatting explicit even when a framework supplies the surrounding retrieval pipeline.