Choose JSON mode, a schema or a tool call
Use structured output when an application needs data with a known shape: order lines, document labels or a short extraction record. A valid JSON object can still contain an invented price or a wrong total. Check structure and business rules before storing or acting on it.
Scroll horizontally to see every column.
| Need | Request | Check after the response |
|---|---|---|
| Parse a JSON reply | JSON mode | Fields, types and values remain application checks |
| Describe an extraction record | JSON Schema structured output | Validate the record, arithmetic and source accuracy |
| Ask an application to perform an action | Tool calling | Validate arguments and authorize the action separately |
Ollama's native chat API accepts a schema through format; its OpenAI-compatible chat API uses response_format. This guide uses /v1/chat/completions. See the Ollama structured-output documentation and the supported OpenAI-compatible fields. For actions, use the tool-calling guide.
This example does not set strict: true or treat a provider's strict-mode flag as an application guarantee. It validates the returned record itself. Temperature 0 does not prove factual accuracy or identical results across runtime versions and hardware.
Start with local Ollama and a synthetic order
Install Node.js 24 and a local Ollama server, then download the test model. Keep Ollama bound to loopback. If you already have a server running, the pull command is enough; do not start a second server.
ollama serveollama pull qwen3:8b
node --versionThe fixture describes order ORD-1042 in EUR: two GPU-CABLE units at 1250 cents each, one FAN-120 at 600 cents, and a 3100-cent total. Integer cents avoid floating-point currency arithmetic. The SKU catalog in the client is trusted application data, separate from the model's reply.
The expected record below is a synthetic test fixture. It is not a measured inference result or evidence that this model reliably extracts arbitrary orders. The current Ollama documentation also states that Ollama Cloud does not support structured outputs; use a local model for this workflow.
Request a schema and validate before producing output
Save this complete client as extract-order.mjs. It uses Node's built-in fetch, sends one non-streaming request and reads the full response under the same deadline. The default is 120 seconds, configurable with LLM_TIMEOUT_MS from 100 to 120000. Redirects are rejected, remote endpoints require HTTPS and a key, and the body is limited to 64 KiB of decoded response bytes. For this Qwen3 example, reasoning_effort: none requests no thinking output on a compatible Ollama runtime.
// Save as extract-order.mjs. Requires Node.js 24; no npm packages.
const SOURCE = "Order ORD-1042 in EUR. GPU-CABLE: 2 units at 1250 cents each. " +
"FAN-120: 1 unit at 600 cents. Total: 3100 cents. No tax or delivery charge.";
const CATALOG = new Map([["GPU-CABLE", 1250], ["FAN-120", 600]]);
const SCHEMA = {
type: "object", additionalProperties: false,
required: ["order_id", "currency", "lines", "total_cents"],
properties: {
order_id: { type: "string", enum: ["ORD-1042"] },
currency: { type: "string", enum: ["EUR"] },
lines: {
type: "array", minItems: 1, maxItems: 2,
items: {
type: "object", additionalProperties: false,
required: ["sku", "quantity", "unit_cents"],
properties: {
sku: { type: "string", enum: [...CATALOG.keys()] },
quantity: { type: "integer", minimum: 1, maximum: 10 },
unit_cents: { type: "integer", minimum: 0, maximum: 1000000 }
}
}
},
total_cents: { type: "integer", minimum: 0, maximum: 20000000 }
}
};
function reject() { throw new Error("invalid"); }
function object(value) {
return value !== null && typeof value === "object" && !Array.isArray(value);
}
function exact(value, fields) {
if (!object(value) || Object.keys(value).length !== fields.length ||
!fields.every((field) => Object.hasOwn(value, field))) reject();
}
function integer(value, min, max) {
if (!Number.isSafeInteger(value) || value < min || value > max) reject();
}
function validate(value) {
exact(value, ["order_id", "currency", "lines", "total_cents"]);
if (value.order_id !== "ORD-1042" || value.currency !== "EUR" ||
!Array.isArray(value.lines) || value.lines.length < 1 ||
value.lines.length > CATALOG.size) reject();
const seen = new Set();
let total = 0;
for (const line of value.lines) {
exact(line, ["sku", "quantity", "unit_cents"]);
if (!CATALOG.has(line.sku) || seen.has(line.sku)) reject();
integer(line.quantity, 1, 10);
integer(line.unit_cents, 0, 1000000);
if (line.unit_cents !== CATALOG.get(line.sku)) reject();
seen.add(line.sku);
total += line.quantity * line.unit_cents;
}
integer(value.total_cents, 0, 20000000);
if (value.total_cents !== total) reject();
return value;
}
async function readJson(response) {
if (!response.body || response.headers.get("content-type")?.split(";")[0]
.trim().toLowerCase() !== "application/json") reject();
const reader = response.body.getReader();
const chunks = [];
let bytes = 0;
try {
for (;;) {
const { done, value } = await reader.read();
if (done) break;
bytes += value.byteLength;
if (bytes > 65536) reject();
chunks.push(value);
}
} finally { reader.releaseLock(); }
return JSON.parse(new TextDecoder("utf-8", { fatal: true })
.decode(Buffer.concat(chunks, bytes)));
}
async function main() {
const baseText = process.env.LLM_BASE_URL ?? "http://127.0.0.1:11434/v1";
if (baseText.length > 2048 || /[\s\u0000-\u001f\u007f]/u.test(baseText)) reject();
const base = new URL(baseText);
const loopback = ["localhost", "127.0.0.1", "[::1]"].includes(base.hostname);
if (base.username || base.password || base.search || base.hash ||
base.pathname.replace(/\/+$/u, "") !== "/v1" ||
(base.protocol !== "https:" && !(base.protocol === "http:" && loopback))) reject();
const key = process.env.LLM_API_KEY ?? "";
if (key.length > 512 || /[^\x21-\x7e]/u.test(key) || (!loopback && !key)) reject();
const model = process.env.LLM_MODEL ?? "qwen3:8b";
if (model.length > 100 || !/^[A-Za-z0-9][A-Za-z0-9._:/-]*$/u.test(model) ||
/[\r\n]/u.test(model)) reject();
const timeoutText = process.env.LLM_TIMEOUT_MS ?? "120000";
const timeoutMs = Number(timeoutText);
if (!/^[0-9]{1,6}$/u.test(timeoutText) || /[\r\n]/u.test(timeoutText) ||
timeoutMs < 100 || timeoutMs > 120000) reject();
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
const response = await fetch(base.origin + "/v1/chat/completions", {
method: "POST", redirect: "error", signal: controller.signal,
headers: {
"Content-Type": "application/json",
...(key ? { Authorization: "Bearer " + key } : {})
},
body: JSON.stringify({
model, stream: false, temperature: 0, max_tokens: 512,
reasoning_effort: "none",
response_format: {
type: "json_schema",
json_schema: { name: "order_summary", schema: SCHEMA }
},
messages: [
{ role: "system", content: "Extract the supplied order. Return only JSON. " +
"Preserve source quantities and amounts. Schema: " + JSON.stringify(SCHEMA) },
{ role: "user", content: SOURCE }
]
})
});
if (!response.ok) reject();
const envelope = await readJson(response);
if (!object(envelope) || envelope.error != null ||
!Array.isArray(envelope.choices) || envelope.choices.length !== 1) reject();
const choice = envelope.choices[0];
if (!object(choice) || choice.index !== 0 || choice.finish_reason !== "stop" ||
!object(choice.message)) reject();
const message = choice.message;
if (message.role !== "assistant" || typeof message.content !== "string" ||
!message.content.trim() || Buffer.byteLength(message.content, "utf8") > 16384 ||
(message.refusal != null && message.refusal !== "") ||
(message.tool_calls != null &&
(!Array.isArray(message.tool_calls) || message.tool_calls.length !== 0))) reject();
const result = validate(JSON.parse(message.content));
process.stdout.write(JSON.stringify(result) + "\n");
} finally {
clearTimeout(timer);
controller.abort();
}
}
main().catch(() => {
// Never print keys, input, raw engine bodies or exception messages.
console.error("Extraction failed: check endpoint, completion and validation.");
process.exitCode = 1;
});The validator rejects extra or missing fields, arrays in place of objects, unknown or repeated SKUs, non-integer quantities and prices that disagree with the trusted catalog. It recomputes the total. It also requires a complete assistant reply and rejects refusals, tool calls and finish_reason: length before parsing the content.
Schema validation alone cannot prove that every line was extracted or that a quantity matches the source. This fixture therefore has an independent expected-result check in the next section. A production extractor needs equivalent source checks or human review for facts that cannot be checked against trusted data.
Verify the fixture, then test failures
Save the verifier below as verify-order.mjs. The command writes an output file and runs the verifier only after the client exits successfully. For this fixture the line order is part of the expected record. A failure exits with a non-zero status and a fixed message; the verifier does not print the rejected record.
// Save as verify-order.mjs beside order.json.
import { readFileSync } from "node:fs";
import { deepStrictEqual } from "node:assert/strict";
try {
const actual = JSON.parse(readFileSync("order.json", "utf8"));
deepStrictEqual(actual, {
"order_id": "ORD-1042",
"currency": "EUR",
"lines": [
{
"sku": "GPU-CABLE",
"quantity": 2,
"unit_cents": 1250
},
{
"sku": "FAN-120",
"quantity": 1,
"unit_cents": 600
}
],
"total_cents": 3100
});
console.log("Synthetic order verified");
} catch {
console.error("Synthetic order verification failed");
process.exitCode = 1;
}node extract-order.mjs > order.json && node verify-order.mjsScroll horizontally to see every column.
| Failure | Client behavior | What to investigate |
|---|---|---|
| HTTP error, redirect or unreadable body | No successful record; fixed error message | Endpoint, authentication, runtime support and response format |
| Deadline or body limit exceeded | Abort the request | Model speed, output length and workload size |
| Truncation, refusal or tool-call reply | Reject before content parsing | Output budget, prompt and chosen model |
| Wrong fields, SKU, price or total | Reject during application validation | Schema, trusted catalog and source facts |
A failed run can leave an empty order.json; downstream code must check the exit status. Do not retry forever or silently accept a partial record. If you add retries, bound their count and total duration, and decide whether an extra inference request is acceptable for the workload.
Use the same client through the gpuOS gateway
Connect your GPU and deploy a model, then select a deployed gpuOS model id. The gateway forwards the request's response_format to the node's Ollama server while mapping that id to its Ollama tag. Runtime and model support still determine the engine's behavior; gpuOS does not add a separate strict-schema validator to your extraction.
# Bash: enter a workspace key without putting it in shell history.
set +x
export LLM_BASE_URL='https://gpuos.si/v1'
export LLM_MODEL='qwen3-8b'
read -r -s -p 'Workspace API key: ' LLM_API_KEY
printf '\n'
export LLM_API_KEY
node extract-order.mjs > order.json && node verify-order.mjs
unset LLM_API_KEYKeep the base URL in trusted application configuration, never in a document or model reply. Inference runs on your connected machine; prompts and outputs pass through the hosted gpuOS gateway. Review that data flow before sending private documents. The client prints fixed errors so an engine error cannot copy a document or key into your logs.
Evaluate extraction quality before a production workflow
- Build a labeled corpus with missing fields, ambiguous quantities, discounts, duplicate lines, multilingual text and instructions embedded in documents. Measure field accuracy separately from JSON validity.
- Keep raw model data out of payment, fulfillment or authorization decisions until application rules and any required review pass. Validate identifiers against the records the caller is allowed to access.
- Record schema and prompt versions, model tag or digest, runtime version, rejection reason and latency. Avoid storing raw private text by default.
- Version currency, tax and discount rules explicitly. This sample only handles EUR integer cents with no tax or delivery charge; changing the domain requires changing the schema, validator and expected fixtures together.
Compare prompt and model changes on the same labeled cases using the local LLM benchmark protocol. Start with the Ollama API guide if the endpoint itself is not yet working.