Open models you can run on your own GPU
A short, tested catalog instead of a model zoo. Each model is pinned to a quantization and an engine (Ollama), with the VRAM it needs at short context and its license. Deploy one in a click and call it through one OpenAI-compatible API.
Chat
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| Qwen3 32B Internal assistants · RAG over company documents · Multilingual chat | 32B dense | ≈ 22 GB | 32K | Apache 2.0 | Pro |
| GLM-4 32B 0414 Tool calling and agents · Structured JSON output · Code generation | 32B dense | ≈ 20.5 GB | 32K | MIT | Pro |
| Mistral Small 3.2 24B European-language chat · Long documents · Instruction following | 24B dense | ≈ 17 GB | 128K | Apache 2.0 | Pro |
| Gemma 3 27B Multilingual chat · Long context · Image understanding | 27B dense | ≈ 18 GB | 128K | Gemma Terms of Use | Pro |
Fast MoE
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| Qwen3 30B A3B High-throughput chat · Latency-sensitive apps · Many concurrent users | 30B MoE (3B active) | ≈ 20 GB | 32K | Apache 2.0 | Pro |
Agentic coding
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| gpt-oss 20B Coding agents · Tool use · Long-context tasks | 21B MoE | ≈ 14 GB | 128K | Apache 2.0 | Free |
Reasoning
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| DeepSeek R1 Distill Qwen 32B Math and logic · Planning · Hard multi-step questions | 32B dense | ≈ 22 GB | 32K | MIT | Pro |
| GLM-Z1 32B 0414 Reasoning · Math · Code review | 32B dense | ≈ 20.5 GB | 32K | MIT | Pro |
Small & fast
Vision
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| Qwen3-VL 8B Document and invoice OCR · Screenshot understanding · Chart reading | 8B dense | ≈ 7.5 GB | 32K | Apache 2.0 | Free |
Embeddings
| Model | Size | VRAM | Context | License | Plan |
|---|---|---|---|---|---|
| BGE-M3 Semantic search · RAG retrieval · Deduplication and clustering | 568M | ≈ 1.5 GB | 8K | MIT | Free |
Run them on your hardware
Connect a GPU with one command, deploy a model, and get an OpenAI-compatible endpoint with keys, quotas and usage metering.