Chat · Pro plan
Run Gemma 3 27B on your own GPU
Google's Gemma 3 27B, a strong multilingual model with a 128K context and image input. Its Gemma Terms of Use add restrictions, so review them before commercial use.
With gpuos, Gemma 3 27B runs on Ollama at Q4_K_M on your machine and is served as gemma-3-27b through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.
- Parameters
- 27B dense
- Quantization
- Q4_K_M
- VRAM (4–8K ctx)
- ≈ 18 GB
- Context window
- 128K tokens
- License
- Gemma Terms of Use
- Engine
- Ollama
- Ollama tag
- gemma3:27b
- API
- /v1/chat/completions
Gemma's terms add use restrictions. Review them before commercial use.
What Gemma 3 27B is good at
- Multilingual chat
- Long context
- Image understanding
Which GPUs can run Gemma 3 27B?
Catalog estimate at Q4_K_M, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.
Check this estimate against your GPU with the VRAM calculator
- RTX 4060 Ti 8 GBToo small
- RTX 4070 Ti Super 16 GBToo small
- RTX 4000 Ada 20 GBFits
- RTX 3090 / 4090 24 GBFits
- RTX PRO 4000 Blackwell 24 GB (Hetzner GEX45)Fits
- NVIDIA L4 24 GBFits
- RTX 5090 32 GBFits
- A100 / H100 80 GBFits
- RTX PRO 6000 Blackwell 96 GB (Hetzner GEX131)Fits
Call Gemma 3 27B with the OpenAI SDK
Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.
from openai import OpenAI
client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
stream = client.chat.completions.create(
model="gemma-3-27b",
messages=[{"role": "user", "content": "Summarize our refund policy."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")import OpenAI from "openai"
const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const reply = await client.chat.completions.create({
model: "gemma-3-27b",
messages: [{ role: "user", content: "Summarize our refund policy." }],
})
console.log(reply.choices[0].message.content)curl https://gpuos.si/v1/chat/completions \
-H "Authorization: Bearer $GPUOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gemma-3-27b", "messages": [{"role": "user", "content": "Hello"}]}'Questions
- How much VRAM does Gemma 3 27B need?
- About 18 GB at Q4_K_M with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 4000 Ada 20 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
- Can I use Gemma 3 27B commercially?
- Gemma 3 27B is released under the Gemma Terms of Use. Gemma's terms add use restrictions. Review them before commercial use.
- Is Gemma 3 27B compatible with the OpenAI API?
- Yes. Through gpuos, Gemma 3 27B is served at /v1/chat/completions with the model id "gemma-3-27b", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
- Which gpuos plan includes Gemma 3 27B?
- Gemma 3 27B is part of the full catalog on the Pro plan, $29 per GPU per month.
Related models
Plan your Gemma 3 27B deployment with gpuOS
Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.