gpuos

Chat · Pro plan

Run Qwen3 32B on your own GPU

Alibaba's dense 32B model, estimated at 22 GB in Q4_K_M with a short context. It supports thinking and non-thinking modes; test answer quality and memory on your own workload.

With gpuos, Qwen3 32B runs on Ollama at Q4_K_M on your machine and is served as qwen3-32b through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.

Parameters
32B dense
Quantization
Q4_K_M
VRAM (4–8K ctx)
≈ 22 GB
Context window
32K tokens
License
Apache 2.0
Engine
Ollama
Ollama tag
qwen3:32b
API
/v1/chat/completions

What Qwen3 32B is good at

Which GPUs can run Qwen3 32B?

Catalog estimate at Q4_K_M, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.

Check this estimate against your GPU with the VRAM calculator

Local setup and workload validation

Install a GPU-compatible Ollama runtime, confirm the driver with nvidia-smi, then pull and inspect the exact tag. Download size is not the same as runtime VRAM.

On your GPU machine
ollama pull qwen3:32b
ollama show qwen3:32b
ollama run qwen3:32b "Reply with a short greeting"
ollama ps

The 22 GB Q4_K_M estimate leaves about 2 GB on a 24 GB card. The advertised context window is not a promise that all its tokens fit on that card. Verify context and concurrency separately.

For a reproducible check, record the GPU, driver, Ollama version, model tag and quantization. Test one short input first, then the actual context and batch/concurrency you need. Separate model loading from warm inference and record peak memory and errors. No throughput benchmark is claimed on this page.

Runtime source: Ollama model card. Tags can change; inspect the pulled model before comparing results.

Call Qwen3 32B with the OpenAI SDK

Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.

Python
from openai import OpenAI

client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
stream = client.chat.completions.create(
    model="qwen3-32b",
    messages=[{"role": "user", "content": "Summarize our refund policy."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Node.js
import OpenAI from "openai"

const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const reply = await client.chat.completions.create({
  model: "qwen3-32b",
  messages: [{ role: "user", content: "Summarize our refund policy." }],
})
console.log(reply.choices[0].message.content)
curl
curl https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-32b", "messages": [{"role": "user", "content": "Hello"}]}'

Questions

How much VRAM does Qwen3 32B need?
About 22 GB at Q4_K_M with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 3090 / 4090 24 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
Can I use Qwen3 32B commercially?
Yes. Qwen3 32B is released under the Apache 2.0 license, which allows commercial use.
Is Qwen3 32B compatible with the OpenAI API?
Yes. Through gpuos, Qwen3 32B is served at /v1/chat/completions with the model id "qwen3-32b", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
Which gpuos plan includes Qwen3 32B?
Qwen3 32B is part of the full catalog on the Pro plan, $29 per GPU per month.

Related models

Plan your Qwen3 32B deployment with gpuOS

Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.