gpuos

Use cases · 8 min read · updated Oct 9, 2026

Ollama structured outputs: JSON Schema and strict client validation

Extract JSON with Ollama's OpenAI-compatible API. Run a Node.js example with JSON Schema, bounded requests and business-rule validation before using the result.

On this page

Choose JSON mode, a schema or a tool call

Use structured output when an application needs data with a known shape: order lines, document labels or a short extraction record. A valid JSON object can still contain an invented price or a wrong total. Check structure and business rules before storing or acting on it.

Scroll horizontally to see every column.

NeedRequestCheck after the response
Parse a JSON replyJSON modeFields, types and values remain application checks
Describe an extraction recordJSON Schema structured outputValidate the record, arithmetic and source accuracy
Ask an application to perform an actionTool callingValidate arguments and authorize the action separately

Ollama's native chat API accepts a schema through format; its OpenAI-compatible chat API uses response_format. This guide uses /v1/chat/completions. See the Ollama structured-output documentation and the supported OpenAI-compatible fields. For actions, use the tool-calling guide.

This example does not set strict: true or treat a provider's strict-mode flag as an application guarantee. It validates the returned record itself. Temperature 0 does not prove factual accuracy or identical results across runtime versions and hardware.

Start with local Ollama and a synthetic order

Install Node.js 24 and a local Ollama server, then download the test model. Keep Ollama bound to loopback. If you already have a server running, the pull command is enough; do not start a second server.

Terminal 1, if Ollama is not running
ollama serve
Terminal 2, prepare the model and check Node.js
ollama pull qwen3:8b
node --version

The fixture describes order ORD-1042 in EUR: two GPU-CABLE units at 1250 cents each, one FAN-120 at 600 cents, and a 3100-cent total. Integer cents avoid floating-point currency arithmetic. The SKU catalog in the client is trusted application data, separate from the model's reply.

The expected record below is a synthetic test fixture. It is not a measured inference result or evidence that this model reliably extracts arbitrary orders. The current Ollama documentation also states that Ollama Cloud does not support structured outputs; use a local model for this workflow.

Request a schema and validate before producing output

Save this complete client as extract-order.mjs. It uses Node's built-in fetch, sends one non-streaming request and reads the full response under the same deadline. The default is 120 seconds, configurable with LLM_TIMEOUT_MS from 100 to 120000. Redirects are rejected, remote endpoints require HTTPS and a key, and the body is limited to 64 KiB of decoded response bytes. For this Qwen3 example, reasoning_effort: none requests no thinking output on a compatible Ollama runtime.

extract-order.mjs, Node.js 24 with no dependencies
// Save as extract-order.mjs. Requires Node.js 24; no npm packages.
const SOURCE = "Order ORD-1042 in EUR. GPU-CABLE: 2 units at 1250 cents each. " +
  "FAN-120: 1 unit at 600 cents. Total: 3100 cents. No tax or delivery charge.";
const CATALOG = new Map([["GPU-CABLE", 1250], ["FAN-120", 600]]);
const SCHEMA = {
  type: "object", additionalProperties: false,
  required: ["order_id", "currency", "lines", "total_cents"],
  properties: {
    order_id: { type: "string", enum: ["ORD-1042"] },
    currency: { type: "string", enum: ["EUR"] },
    lines: {
      type: "array", minItems: 1, maxItems: 2,
      items: {
        type: "object", additionalProperties: false,
        required: ["sku", "quantity", "unit_cents"],
        properties: {
          sku: { type: "string", enum: [...CATALOG.keys()] },
          quantity: { type: "integer", minimum: 1, maximum: 10 },
          unit_cents: { type: "integer", minimum: 0, maximum: 1000000 }
        }
      }
    },
    total_cents: { type: "integer", minimum: 0, maximum: 20000000 }
  }
};
function reject() { throw new Error("invalid"); }
function object(value) {
  return value !== null && typeof value === "object" && !Array.isArray(value);
}
function exact(value, fields) {
  if (!object(value) || Object.keys(value).length !== fields.length ||
      !fields.every((field) => Object.hasOwn(value, field))) reject();
}
function integer(value, min, max) {
  if (!Number.isSafeInteger(value) || value < min || value > max) reject();
}
function validate(value) {
  exact(value, ["order_id", "currency", "lines", "total_cents"]);
  if (value.order_id !== "ORD-1042" || value.currency !== "EUR" ||
      !Array.isArray(value.lines) || value.lines.length < 1 ||
      value.lines.length > CATALOG.size) reject();
  const seen = new Set();
  let total = 0;
  for (const line of value.lines) {
    exact(line, ["sku", "quantity", "unit_cents"]);
    if (!CATALOG.has(line.sku) || seen.has(line.sku)) reject();
    integer(line.quantity, 1, 10);
    integer(line.unit_cents, 0, 1000000);
    if (line.unit_cents !== CATALOG.get(line.sku)) reject();
    seen.add(line.sku);
    total += line.quantity * line.unit_cents;
  }
  integer(value.total_cents, 0, 20000000);
  if (value.total_cents !== total) reject();
  return value;
}
async function readJson(response) {
  if (!response.body || response.headers.get("content-type")?.split(";")[0]
      .trim().toLowerCase() !== "application/json") reject();
  const reader = response.body.getReader();
  const chunks = [];
  let bytes = 0;
  try {
    for (;;) {
      const { done, value } = await reader.read();
      if (done) break;
      bytes += value.byteLength;
      if (bytes > 65536) reject();
      chunks.push(value);
    }
  } finally { reader.releaseLock(); }
  return JSON.parse(new TextDecoder("utf-8", { fatal: true })
    .decode(Buffer.concat(chunks, bytes)));
}
async function main() {
  const baseText = process.env.LLM_BASE_URL ?? "http://127.0.0.1:11434/v1";
  if (baseText.length > 2048 || /[\s\u0000-\u001f\u007f]/u.test(baseText)) reject();
  const base = new URL(baseText);
  const loopback = ["localhost", "127.0.0.1", "[::1]"].includes(base.hostname);
  if (base.username || base.password || base.search || base.hash ||
      base.pathname.replace(/\/+$/u, "") !== "/v1" ||
      (base.protocol !== "https:" && !(base.protocol === "http:" && loopback))) reject();
  const key = process.env.LLM_API_KEY ?? "";
  if (key.length > 512 || /[^\x21-\x7e]/u.test(key) || (!loopback && !key)) reject();
  const model = process.env.LLM_MODEL ?? "qwen3:8b";
  if (model.length > 100 || !/^[A-Za-z0-9][A-Za-z0-9._:/-]*$/u.test(model) ||
      /[\r\n]/u.test(model)) reject();
  const timeoutText = process.env.LLM_TIMEOUT_MS ?? "120000";
  const timeoutMs = Number(timeoutText);
  if (!/^[0-9]{1,6}$/u.test(timeoutText) || /[\r\n]/u.test(timeoutText) ||
      timeoutMs < 100 || timeoutMs > 120000) reject();
  const controller = new AbortController();
  const timer = setTimeout(() => controller.abort(), timeoutMs);
  try {
    const response = await fetch(base.origin + "/v1/chat/completions", {
      method: "POST", redirect: "error", signal: controller.signal,
      headers: {
        "Content-Type": "application/json",
        ...(key ? { Authorization: "Bearer " + key } : {})
      },
      body: JSON.stringify({
        model, stream: false, temperature: 0, max_tokens: 512,
        reasoning_effort: "none",
        response_format: {
          type: "json_schema",
          json_schema: { name: "order_summary", schema: SCHEMA }
        },
        messages: [
          { role: "system", content: "Extract the supplied order. Return only JSON. " +
            "Preserve source quantities and amounts. Schema: " + JSON.stringify(SCHEMA) },
          { role: "user", content: SOURCE }
        ]
      })
    });
    if (!response.ok) reject();
    const envelope = await readJson(response);
    if (!object(envelope) || envelope.error != null ||
        !Array.isArray(envelope.choices) || envelope.choices.length !== 1) reject();
    const choice = envelope.choices[0];
    if (!object(choice) || choice.index !== 0 || choice.finish_reason !== "stop" ||
        !object(choice.message)) reject();
    const message = choice.message;
    if (message.role !== "assistant" || typeof message.content !== "string" ||
        !message.content.trim() || Buffer.byteLength(message.content, "utf8") > 16384 ||
        (message.refusal != null && message.refusal !== "") ||
        (message.tool_calls != null &&
          (!Array.isArray(message.tool_calls) || message.tool_calls.length !== 0))) reject();
    const result = validate(JSON.parse(message.content));
    process.stdout.write(JSON.stringify(result) + "\n");
  } finally {
    clearTimeout(timer);
    controller.abort();
  }
}
main().catch(() => {
  // Never print keys, input, raw engine bodies or exception messages.
  console.error("Extraction failed: check endpoint, completion and validation.");
  process.exitCode = 1;
});

The validator rejects extra or missing fields, arrays in place of objects, unknown or repeated SKUs, non-integer quantities and prices that disagree with the trusted catalog. It recomputes the total. It also requires a complete assistant reply and rejects refusals, tool calls and finish_reason: length before parsing the content.

Schema validation alone cannot prove that every line was extracted or that a quantity matches the source. This fixture therefore has an independent expected-result check in the next section. A production extractor needs equivalent source checks or human review for facts that cannot be checked against trusted data.

Verify the fixture, then test failures

Save the verifier below as verify-order.mjs. The command writes an output file and runs the verifier only after the client exits successfully. For this fixture the line order is part of the expected record. A failure exits with a non-zero status and a fixed message; the verifier does not print the rejected record.

verify-order.mjs, compare with the known synthetic order
// Save as verify-order.mjs beside order.json.
import { readFileSync } from "node:fs";
import { deepStrictEqual } from "node:assert/strict";
try {
  const actual = JSON.parse(readFileSync("order.json", "utf8"));
  deepStrictEqual(actual, {
  "order_id": "ORD-1042",
  "currency": "EUR",
  "lines": [
    {
      "sku": "GPU-CABLE",
      "quantity": 2,
      "unit_cents": 1250
    },
    {
      "sku": "FAN-120",
      "quantity": 1,
      "unit_cents": 600
    }
  ],
  "total_cents": 3100
});
  console.log("Synthetic order verified");
} catch {
  console.error("Synthetic order verification failed");
  process.exitCode = 1;
}
Run locally and verify the result
node extract-order.mjs > order.json && node verify-order.mjs

Scroll horizontally to see every column.

FailureClient behaviorWhat to investigate
HTTP error, redirect or unreadable bodyNo successful record; fixed error messageEndpoint, authentication, runtime support and response format
Deadline or body limit exceededAbort the requestModel speed, output length and workload size
Truncation, refusal or tool-call replyReject before content parsingOutput budget, prompt and chosen model
Wrong fields, SKU, price or totalReject during application validationSchema, trusted catalog and source facts

A failed run can leave an empty order.json; downstream code must check the exit status. Do not retry forever or silently accept a partial record. If you add retries, bound their count and total duration, and decide whether an extra inference request is acceptable for the workload.

Use the same client through the gpuOS gateway

Connect your GPU and deploy a model, then select a deployed gpuOS model id. The gateway forwards the request's response_format to the node's Ollama server while mapping that id to its Ollama tag. Runtime and model support still determine the engine's behavior; gpuOS does not add a separate strict-schema validator to your extraction.

Call your deployed gpuOS model with a workspace API key
# Bash: enter a workspace key without putting it in shell history.
set +x
export LLM_BASE_URL='https://gpuos.si/v1'
export LLM_MODEL='qwen3-8b'
read -r -s -p 'Workspace API key: ' LLM_API_KEY
printf '\n'
export LLM_API_KEY
node extract-order.mjs > order.json && node verify-order.mjs
unset LLM_API_KEY

Keep the base URL in trusted application configuration, never in a document or model reply. Inference runs on your connected machine; prompts and outputs pass through the hosted gpuOS gateway. Review that data flow before sending private documents. The client prints fixed errors so an engine error cannot copy a document or key into your logs.

Evaluate extraction quality before a production workflow

  • Build a labeled corpus with missing fields, ambiguous quantities, discounts, duplicate lines, multilingual text and instructions embedded in documents. Measure field accuracy separately from JSON validity.
  • Keep raw model data out of payment, fulfillment or authorization decisions until application rules and any required review pass. Validate identifiers against the records the caller is allowed to access.
  • Record schema and prompt versions, model tag or digest, runtime version, rejection reason and latency. Avoid storing raw private text by default.
  • Version currency, tax and discount rules explicitly. This sample only handles EUR integer cents with no tax or delivery charge; changing the domain requires changing the schema, validator and expected fixtures together.

Compare prompt and model changes on the same labeled cases using the local LLM benchmark protocol. Start with the Ollama API guide if the endpoint itself is not yet working.

Questions

What is the difference between Ollama JSON mode and structured outputs?
JSON mode requests JSON. Structured outputs attach a JSON Schema describing the record. The application still needs to validate fields and business rules and check facts against the source.
How do I request structured output through the OpenAI-compatible API?
Use response_format with type json_schema on /v1/chat/completions. The native Ollama chat API instead uses format. Test the installed runtime and model with your schema before relying on the result.
Does temperature 0 make extraction deterministic and accurate?
No. It changes sampling, but it does not establish factual accuracy or identical results across runtimes and hardware. Use labeled fixtures, application validation and source checks.
Can the same example run through gpuOS?
Yes, configure https://gpuos.si/v1, a workspace API key and a deployed gpuOS model id. The request goes through the hosted gateway to Ollama on your connected node. Engine support and application validation remain necessary.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.