gpuos

Use cases · 8 min read · updated Oct 7, 2026

Self-hosted embeddings API: BGE-M3 vs Qwen3 Embedding

Build a runnable embeddings API retrieval example with BGE-M3 or Qwen3 Embedding. Compare memory estimates, query instructions and index migration.

On this page

Choose an embedding model for retrieval, not chat

An embeddings API turns text into vectors you can compare for semantic retrieval. BGE-M3 and Qwen3 Embedding 0.6B are two catalog candidates for this stage of a RAG pipeline. They do not generate the final answer; use a chat model after retrieving relevant passages.

This tutorial builds a small, runnable dense-retrieval client with Python's standard library. Run it against local Ollama first, then use the same request shape with a deployed gpuOS model. The fixture demonstrates API handling and ranking; it is not a benchmark of either model.

Scroll horizontally to see every column.

ModelLocal Ollama tagCatalog quantizationEstimated VRAMCatalog context (tokens)License
BGE-M3bge-m3FP161.5 GB8,192MIT
Qwen3 Embedding 0.6Bqwen3-embedding:0.6bQ8_02 GB32,768Apache 2.0

Memory values are planning estimates for short inputs. Long chunks and large batches can raise peak memory. The context column is an input limit, not a recommended chunk size or a promise that a maximum-size batch fits.

The differences that change your client and index

Scroll horizontally to see every column.

DecisionBGE-M3Qwen3 Embedding 0.6B
Query formattingNo retrieval instruction requiredPrefix queries with a task instruction; leave documents unprefixed
Publisher vector dimension1024Up to 1024; the publisher supports smaller dimensions
API recipe hereOne dense vector per textOne dense vector per text, default dimensions
Switching modelsRebuild document vectors and re-evaluate retrievalRebuild document vectors and re-evaluate retrieval

The BGE-M3 publisher also provides sparse and multi-vector retrieval through its own tooling. This OpenAI-compatible embeddings recipe uses dense vectors only; selecting BGE-M3 does not automatically create a hybrid retriever. BAAI model card.

Qwen's query instruction describes the retrieval task; it is not an instruction for a chat assistant. The script adds it only to the query. It leaves output dimensions at the server default and reads the actual response size rather than assuming a custom-dimension API works in every runtime. Qwen model card.

Start an API on the machine running Ollama

Prerequisites: Python 3, an installed and running Ollama server, enough RAM/VRAM for your selected model, and disk space for its weights. Keep the default local bind for this test. Pull only the candidate you plan to evaluate first.

Check the OpenAI-compatible embeddings route
ollama --version
ollama pull bge-m3
curl --fail-with-body http://127.0.0.1:11434/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-m3","input":["Rotate a leaked API key.","VRAM stores GPU model weights."],"encoding_format":"float"}'

Expect a JSON data array with one indexed embedding vector per input. Ollama's native /api/embed returns a different shape with an embeddings array. Choose one protocol for your client; the script below uses /v1/embeddings. Ollama OpenAI compatibility.

For Qwen, pull qwen3-embedding:0.6b and pass that exact local tag. A connection refusal means the local server is unavailable; a model-not-found response means the selected tag has not been downloaded or does not match.

Run a complete dense-retrieval example

Save the following as retrieve.py. It embeds three short passages and one question, restores the API's indexed order, checks numeric vectors and computes cosine similarity. It runs with no pip dependencies. The client refuses redirects; configure the final API URL directly so the key stays on the intended endpoint.

retrieve.py
import json
import math
import os
import urllib.error
import urllib.request

BASE = os.environ.get("EMBEDDING_BASE_URL", "http://127.0.0.1:11434/v1").rstrip("/")
MODEL = os.environ.get("EMBEDDING_MODEL", "bge-m3")
KEY = os.environ.get("GPUOS_API_KEY", "")
DOCUMENTS = [
    ("gpu", "The GPU node runs the model. Its VRAM must fit weights and the workload."),
    ("key", "An API key authorizes access. Revoke a leaked key and issue a replacement."),
    ("index", "An embedding index stores document vectors for semantic retrieval."),
]
QUESTION = "What should I do if an API key leaks?"


class NoRedirectHandler(urllib.request.HTTPRedirectHandler):
    def redirect_request(self, request, fp, code, message, headers, new_url):
        return None


# Refuse redirects so an Authorization header cannot reach a different endpoint.
CLIENT = urllib.request.build_opener(NoRedirectHandler())


def query_text(question):
    if MODEL.startswith("qwen3-embedding"):
        return "Instruct: Retrieve the passage that answers the question.\nQuery:" + question
    return question


def embed(texts):
    if any(ord(character) <= 32 or ord(character) >= 127 for character in KEY):
        raise ValueError("API key must contain printable ASCII characters without spaces")
    headers = {"Content-Type": "application/json"}
    if KEY:
        headers["Authorization"] = "Bearer " + KEY
    request = urllib.request.Request(
        BASE + "/embeddings",
        data=json.dumps({"model": MODEL, "input": texts, "encoding_format": "float"}).encode(),
        headers=headers,
        method="POST",
    )
    with CLIENT.open(request, timeout=120) as response:
        payload = json.load(response)
    if not isinstance(payload, dict):
        raise ValueError("Expected an embedding response object")
    rows = payload.get("data")
    if not isinstance(rows, list) or len(rows) != len(texts):
        raise ValueError("Expected one embedding row per input")
    by_index = {}
    for row in rows:
        if not isinstance(row, dict):
            raise ValueError("Expected an embedding row object")
        index = row.get("index")
        vector = row.get("embedding")
        if type(index) is not int or not 0 <= index < len(texts) or index in by_index:
            raise ValueError("Invalid or duplicate embedding index")
        if not isinstance(vector, list) or not vector or any(
            type(value) not in (int, float) or not math.isfinite(value) for value in vector
        ):
            raise ValueError("Expected a nonempty finite numeric vector")
        by_index[index] = vector
    vectors = [by_index[index] for index in range(len(texts))]
    if len({len(vector) for vector in vectors}) != 1:
        raise ValueError("Embedding dimensions differ within the batch")
    return vectors


def cosine(left, right):
    if len(left) != len(right):
        raise ValueError("Embedding dimensions differ")
    denominator = math.sqrt(sum(x * x for x in left) * sum(x * x for x in right))
    if denominator == 0:
        raise ValueError("Cannot rank a zero vector")
    return sum(a * b for a, b in zip(left, right)) / denominator


def main():
    vectors = embed([text for _, text in DOCUMENTS] + [query_text(QUESTION)])
    ranking = sorted(
        ((cosine(vector, vectors[-1]), doc_id, text)
         for (doc_id, text), vector in zip(DOCUMENTS, vectors[:-1])),
        reverse=True,
    )
    print("model=" + MODEL, "dimensions=" + str(len(vectors[0])))
    for score, doc_id, text in ranking:
        print(f"{score:.4f} {doc_id}: {text}")


if __name__ == "__main__":
    try:
        main()
    except urllib.error.HTTPError as error:
        raise SystemExit(f"Embedding HTTP {error.code}: check model, key and server logs")
    except (urllib.error.URLError, TimeoutError, ValueError, KeyError, TypeError) as error:
        raise SystemExit(f"Embedding request failed: {error}")
Run both local candidates
python3 retrieve.py

# Compare the second candidate with the same document fixture.
ollama pull qwen3-embedding:0.6b
EMBEDDING_MODEL=qwen3-embedding:0.6b python3 retrieve.py

The output prints the model, actual dimension and three ranked passages. For this question, the key passage is the relevant item to check. Scores and ordering depend on the model and runtime; there is no fixed expected score. A valid response with poor ranking is a retrieval-quality issue, not an HTTP success criterion.

This tiny fixture is an API smoke test. For a real index, embed documents during ingestion, store their vectors and metadata, then embed only the user's query during retrieval. Do not re-embed the whole corpus on every request.

Use the same client with a deployed gpuOS model

Connect a GPU node through the quickstart, deploy BGE-M3 or Qwen3 Embedding 0.6B, and create a server-side workspace API key. Both are currently Community catalog models. Set GPUOS_API_KEY through your environment or secret manager before running this command.

Switch the endpoint and model ID
# GPUOS_API_KEY must already be set in the environment.
EMBEDDING_BASE_URL=https://gpuos.si/v1 \
EMBEDDING_MODEL=bge-m3 \
python3 retrieve.py

# Deploy Qwen3 Embedding first, then use its catalog ID here.
EMBEDDING_BASE_URL=https://gpuos.si/v1 \
EMBEDDING_MODEL=qwen3-embedding-0.6b \
python3 retrieve.py

gpuOS accepts catalog IDs such as qwen3-embedding-0.6b; local Ollama uses tags such as qwen3-embedding:0.6b. The gateway forwards embeddings requests to the connected node's Ollama route. Your input passages and returned vectors pass through the hosted gpuOS gateway, which authorizes requests and meters usage. This is not an entirely local or air-gapped request path.

  • 401: check the workspace key and Bearer header; keep keys out of browser bundles.
  • 404: check the gpuOS catalog ID rather than the Ollama tag.
  • 429: inspect the key's request limit or monthly token quota before retrying.
  • 503: confirm that the model is deployed and its node is connected.
  • 400 or 502 with an engine error: inspect the upstream input/batch limits and runtime logs; reduce the batch or chunk size before trying again.

Compare retrieval quality, latency and memory

  1. Build a labeled set from actual queries, including abbreviations, multilingual inputs and queries with no answer in the corpus. Keep the same documents and chunk boundaries for both candidates.
  2. Measure whether an expected passage appears in the first k results (recall at k). Inspect incorrect top results and near-duplicate chunks instead of comparing raw cosine scores across models.
  3. Record query latency separately from ingestion throughput. Measure cold model loading and warm requests separately, and vary batch size on the hardware you will use.
  4. Record peak memory while embedding realistic chunks. If a maximum-size input fails or truncates, shorten the chunks and re-evaluate answer coverage.
  5. Store the model/tag or digest, quantization, actual vector dimension, query-instruction template and chunking version with the index. Keep this configuration fixed during evaluation.

Embedding scores are ranking values, not calibrated probabilities. Set any relevance threshold using your own positive and negative examples. For RAG, retrieve the passages first and pass their text plus source IDs to a chat model using the private RAG guide.

When testing the native Ollama embedding API, truncate:false lets you reject inputs above its context limit instead of silently shortening them. That is an /api/embed option; do not assume the OpenAI-compatible endpoint accepts the same option. Native embedding request options.

Change models without corrupting the index

The same vector dimension does not make two embedding spaces interchangeable. If you move from BGE-M3 to Qwen3 Embedding, build a new document index, evaluate it with Qwen-formatted queries, then switch both ingestion and retrieval together. Keep the previous index until you can roll back both sides.

Never mix old document vectors with new query vectors or append a new model's vectors into an existing collection. For instruction or chunking changes, keep a separate evaluation/version as well. Plan the backfill time, disk space, access-control metadata and deletion behavior before moving a production corpus.

Use LlamaIndex or LangChain when your application needs document loaders and a persistent vector store. Keep the embedding model and query formatting explicit even when a framework supplies the surrounding retrieval pipeline.

Questions

Is BGE-M3 or Qwen3 Embedding better for my RAG application?
Compare both on labeled queries from your own corpus with identical chunks. Check recall at k, failures by language, query latency, ingestion throughput and memory. A publisher leaderboard cannot determine the winner for your documents.
Should I add an instruction to both queries and documents?
In this retrieval recipe, Qwen queries receive a task instruction and documents remain unprefixed. BGE-M3 queries do not require that instruction. Keep preprocessing consistent between evaluation and production.
Can I replace BGE-M3 with Qwen without rebuilding vectors?
No. Matching dimensions do not mean matching embedding spaces. Re-embed the documents into a separate index and switch the query model with it after evaluation.
Does the BGE-M3 embeddings endpoint provide hybrid search?
This API recipe returns dense vectors. BGE-M3's publisher also offers sparse and multi-vector representations in its own tooling, but those do not become a hybrid search pipeline through this request automatically.
Does using my GPU keep every embedding request on my machine?
A direct local Ollama request runs on the local endpoint. A gpuOS request passes its input and output through the hosted gpuOS gateway to your node. Review the data path and deployment configuration before indexing sensitive documents.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.