gpuos

Chat · Pro plan

Run Mistral Small 3.2 24B on your own GPU

Mistral Small 3.2 24B, a European dense model with a 128K context window, good instruction following and vision input, under Apache 2.0.

With gpuos, Mistral Small 3.2 24B runs on Ollama at Q4_K_M on your machine and is served as mistral-small-24b through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.

Parameters
24B dense
Quantization
Q4_K_M
VRAM (4–8K ctx)
≈ 17 GB
Context window
128K tokens
License
Apache 2.0
Engine
Ollama
Ollama tag
mistral-small3.2:24b
API
/v1/chat/completions

What Mistral Small 3.2 24B is good at

Which GPUs can run Mistral Small 3.2 24B?

Catalog estimate at Q4_K_M, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.

Check this estimate against your GPU with the VRAM calculator

Call Mistral Small 3.2 24B with the OpenAI SDK

Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.

Python
from openai import OpenAI

client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
stream = client.chat.completions.create(
    model="mistral-small-24b",
    messages=[{"role": "user", "content": "Summarize our refund policy."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Node.js
import OpenAI from "openai"

const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const reply = await client.chat.completions.create({
  model: "mistral-small-24b",
  messages: [{ role: "user", content: "Summarize our refund policy." }],
})
console.log(reply.choices[0].message.content)
curl
curl https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "mistral-small-24b", "messages": [{"role": "user", "content": "Hello"}]}'

Questions

How much VRAM does Mistral Small 3.2 24B need?
About 17 GB at Q4_K_M with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 4000 Ada 20 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
Can I use Mistral Small 3.2 24B commercially?
Yes. Mistral Small 3.2 24B is released under the Apache 2.0 license, which allows commercial use.
Is Mistral Small 3.2 24B compatible with the OpenAI API?
Yes. Through gpuos, Mistral Small 3.2 24B is served at /v1/chat/completions with the model id "mistral-small-24b", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
Which gpuos plan includes Mistral Small 3.2 24B?
Mistral Small 3.2 24B is part of the full catalog on the Pro plan, $29 per GPU per month.

Related models

Plan your Mistral Small 3.2 24B deployment with gpuOS

Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.