gpuos

Use cases · 4 min read · updated Oct 7, 2026

LLM tool calling with a self-hosted OpenAI-compatible API

Implement function calling with a local LLM through chat completions. Define tools, validate arguments, execute application code and return tool results.

On this page

The model proposes a function call; your application executes it

Tool calling lets a model propose a named function and structured arguments, then use the returned result in an answer. Your application supplies the functions and controls their execution. Ollama's tool calling documentation shows this request, execution and follow-up pattern for supported models.

A customer-support assistant might request an order lookup; a document assistant might request a search. In both cases, the application checks what the signed-in user can access before returning data. The model does not acquire database credentials or permission to run arbitrary code because a function name appears in its response.

Check the model and chat completions route before integrating

gpuOS provides OpenAI-compatible chat completions, completions and embeddings. This guide uses /v1/chat/completions for tool requests. An OpenAI-compatible base URL does not imply that every OpenAI product feature or endpoint is available. Confirm the deployed model's tool capability and test the exact fields your client sends. Ollama lists its compatibility fields.

Start with a non-streaming request, one small function and a harmless input. This makes the response easy to inspect before you add streaming, multiple functions or a longer conversation. Use the catalog id of the model deployed in gpuOS, rather than assuming the local Ollama tag and public API id are identical.

Inference runs in Ollama on your connected GPU. Requests and outputs pass through the hosted gpuOS gateway. Treat tool schemas and returned data as part of that request flow; return only the application fields needed to complete the current task.

Describe a narrow, useful function with a JSON schema

The example below describes a lookup of a synthetic public product SKU. Run it after deploying qwen3-8b and exporting a workspace key as GPUOS_API_KEY. It sends the first stage of the flow only: the response is a proposed call or an ordinary answer. It neither queries a database nor completes the follow-up on its own.

Request a proposed tool call through gpuOS
curl --fail-with-body https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-8b",
    "stream": false,
    "messages": [{"role": "user", "content": "Look up the public product DEMO-001."}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "lookup_public_product",
        "description": "Look up a public product by its exact SKU.",
        "parameters": {
          "type": "object",
          "properties": {"sku": {"type": "string"}},
          "required": ["sku"],
          "additionalProperties": false
        }
      }
    }]
  }'

Give each function one responsibility and describe when it is useful. Prefer an explicit identifier over a vague string such as ‘search anything’. Keep required fields small and stable. A schema guides generation; your application still validates the decoded arguments before calling its implementation.

Validate the call, return its result and make the next request

Scroll horizontally to see every column.

StageApplication behavior
Read the responseHandle ordinary content, no tool call or one or more proposed calls
Resolve the functionMatch its name against a fixed allowlist
Validate argumentsParse JSON and check allowed fields, types and value limits
Check accessApply the signed-in user's permissions in application code
Execute and replyReturn a bounded result linked to the matching tool call id
Continue the conversationInclude the assistant call and tool result in the next chat request

For the OpenAI chat format, preserve the assistant message containing tool_calls, then add a tool message with the matching tool_call_id and a string result. Keep the prior conversation when making the follow-up. Handle a function error as a clear application outcome so the model can explain it, instead of inventing a successful lookup.

Bound the number of tool rounds and the duration of each function. A model can request another lookup, return incomplete arguments or repeat a previous call. Stop the loop with an explicit outcome when its budget is reached. A client-side timeout alone does not define what happens to a function already running in your application.

Separate read operations from changes that need confirmation

A catalog lookup and a payment have different consequences. For operations that change application state, decide the confirmation and authorization rules in product code. A user should see the exact proposed change before approving it when your workflow requires approval. Keep that policy consistent even if the model's wording changes.

Use an idempotency mechanism for operations that must not be repeated after a retry. Record the application action identifier separately from the conversation. Restrict tool results to useful fields, and treat retrieved text as data rather than new instructions. These choices make tool execution easier to audit and keep the function contract understandable to its users.

Test function selection and end-to-end outcomes separately

  • A request that needs the function, with a valid known identifier.
  • A request that can be answered without calling any function.
  • A missing identifier, malformed argument or unknown function name.
  • A lookup that returns no result, a timeout or an access denial.
  • A repeated call and a conversation that reaches the configured tool-round limit.

Score the selected function, argument validity, permission handling and final answer independently. This distinguishes model failures from dispatcher bugs. Measure the complete flow as well: multiple inference requests plus the tool's own latency can dominate the user-visible wait. Re-run these cases after changing the model, precision, prompt or client library.

Questions

Does a local LLM execute Python functions automatically?
No. It proposes a function name and arguments. Your application validates them, checks access, runs an allowed implementation and returns the result in a follow-up chat request.
Can any OpenAI-compatible model call tools?
Tool behavior depends on the model and serving implementation. Check the deployed model's capability and test its requests and responses through the chat completions endpoint before relying on it.
Can I use a tool call to authorize a database update?
Treat the call as a proposal. Application code must verify the user's permission, validate the change and apply any required confirmation and retry controls before writing data.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.