gpuos

Open models you can run on your own GPU

A short, tested catalog instead of a model zoo. Each model is pinned to a quantization and an engine (Ollama), with the VRAM it needs at short context and its license. Deploy one in a click and call it through one OpenAI-compatible API.

Chat

ModelSizeVRAMContextLicensePlan
Qwen3 32B
Internal assistants · RAG over company documents · Multilingual chat
32B dense≈ 22 GB32KApache 2.0Pro
GLM-4 32B 0414
Tool calling and agents · Structured JSON output · Code generation
32B dense≈ 20.5 GB32KMITPro
Mistral Small 3.2 24B
European-language chat · Long documents · Instruction following
24B dense≈ 17 GB128KApache 2.0Pro
Gemma 3 27B
Multilingual chat · Long context · Image understanding
27B dense≈ 18 GB128KGemma Terms of UsePro

Fast MoE

ModelSizeVRAMContextLicensePlan
Qwen3 30B A3B
High-throughput chat · Latency-sensitive apps · Many concurrent users
30B MoE (3B active)≈ 20 GB32KApache 2.0Pro

Agentic coding

ModelSizeVRAMContextLicensePlan
gpt-oss 20B
Coding agents · Tool use · Long-context tasks
21B MoE≈ 14 GB128KApache 2.0Free

Reasoning

ModelSizeVRAMContextLicensePlan
DeepSeek R1 Distill Qwen 32B
Math and logic · Planning · Hard multi-step questions
32B dense≈ 22 GB32KMITPro
GLM-Z1 32B 0414
Reasoning · Math · Code review
32B dense≈ 20.5 GB32KMITPro

Small & fast

ModelSizeVRAMContextLicensePlan
Qwen3 8B
Classification and extraction · High-volume tasks · Small GPUs
8B dense≈ 6.5 GB32KApache 2.0Free
GLM-4 9B
Long context on small GPUs · Multilingual chat · Summaries
9B dense≈ 6.5 GB128KMITFree

Vision

ModelSizeVRAMContextLicensePlan
Qwen3-VL 8B
Document and invoice OCR · Screenshot understanding · Chart reading
8B dense≈ 7.5 GB32KApache 2.0Free

Embeddings

ModelSizeVRAMContextLicensePlan
BGE-M3
Semantic search · RAG retrieval · Deduplication and clustering
568M≈ 1.5 GB8KMITFree

Run them on your hardware

Connect a GPU with one command, deploy a model, and get an OpenAI-compatible endpoint with keys, quotas and usage metering.