gpuos

Embeddings · Free plan

Run BGE-M3 on your own GPU

BGE-M3 is a multilingual embedding model for semantic search and RAG. It covers 100+ languages and inputs up to 8K tokens, and needs under 2 GB of VRAM.

With gpuos, BGE-M3 runs on Ollama at FP16 on your machine and is served as bge-m3 through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.

Parameters
568M
Quantization
FP16
VRAM (4–8K ctx)
≈ 1.5 GB
Context window
8K tokens
License
MIT
Engine
Ollama
Ollama tag
bge-m3
API
/v1/embeddings

What BGE-M3 is good at

Which GPUs can run BGE-M3?

Catalog estimate at FP16, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.

Check this estimate against your GPU with the VRAM calculator

Local setup and workload validation

Install a GPU-compatible Ollama runtime, confirm the driver with nvidia-smi, then pull and inspect the exact tag. Download size is not the same as runtime VRAM.

On your GPU machine
ollama pull bge-m3
ollama show bge-m3
curl http://localhost:11434/api/embed -H "Content-Type: application/json" -d '{"model":"bge-m3","input":"A short retrieval test"}'
ollama ps

BGE-M3 is an embedding model, not a chat model. The 1.5 GB FP16 estimate is a planning figure; batch size and input length affect memory. Use the embeddings endpoint and keep the same model for indexing and retrieval.

For a reproducible check, record the GPU, driver, Ollama version, model tag and quantization. Test one short input first, then the actual context and batch/concurrency you need. Separate model loading from warm inference and record peak memory and errors. No throughput benchmark is claimed on this page.

Runtime source: Ollama model card. Tags can change; inspect the pulled model before comparing results.

Call BGE-M3 with the OpenAI SDK

Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.

Python
from openai import OpenAI

client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
result = client.embeddings.create(model="bge-m3", input=["first text", "second text"])
print(len(result.data[0].embedding))
Node.js
import OpenAI from "openai"

const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const result = await client.embeddings.create({ model: "bge-m3", input: "first text" })
curl
curl https://gpuos.si/v1/embeddings \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "bge-m3", "input": "first text"}'

Questions

How much VRAM does BGE-M3 need?
About 1.5 GB at FP16 with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 4060 Ti 8 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
Can I use BGE-M3 commercially?
Yes. BGE-M3 is released under the MIT license, which allows commercial use.
Is BGE-M3 compatible with the OpenAI API?
Yes. Through gpuos, BGE-M3 is served at /v1/embeddings with the model id "bge-m3", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
Which gpuos plan includes BGE-M3?
BGE-M3 is in the base catalog, available on the free Community plan (1 node, 1 GPU).

Related models

Plan your BGE-M3 deployment with gpuOS

Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.