gpuos

Use cases · 5 min read · updated Oct 7, 2026

Set up a local AI coding assistant with Ollama and Continue

Configure Continue with a local Ollama coding model. Control repository context, review small edits and verify changes with your project's tests.

On this page

Start with a reviewed edit in an ordinary repository

A local AI coding assistant combines an editor, repository context and a locally served model. Start with a small edit whose behavior you can verify: explain a function, add an input check or repair one failing test. Keep your existing version-control and review workflow. A model response is a proposed change; your project's tests and human review establish whether that change belongs in the codebase.

This setup uses Continue with a direct local Ollama endpoint. It demonstrates chat and editing, without assuming that the chosen model supports autonomous tool use. The local AI guide explains the component choices. If strict offline operation matters, also complete the offline assistant verification for the editor, server and any tools.

Install the interface and pull the exact coding model

Prepare the verified example tag on your local Ollama server
ollama --version
ollama pull qwen2.5-coder:7b
ollama list

Install Continue in your editor and use the runtime's supported installation procedure. The Ollama model listing verifies this tag, while Continue's Ollama guide explains that a model configuration does not download its weights. Check the exact installed name rather than mixing a catalog ID, display name and runtime tag.

This model is an example, not an assertion that it is the best available coding model or fits every laptop. Choose a smaller supported variant when necessary and measure your intended prompts. Coding requests include file excerpts, instructions and an output allowance, so memory and latency should be checked with that actual context rather than a short greeting.

Use the current YAML provider configuration

A minimal local Continue config.yaml for chat, edit and apply roles
name: Local coding review
version: 1.0.0
schema: v1
models:
  - name: Local Qwen coder
    provider: ollama
    model: qwen2.5-coder:7b
    apiBase: http://127.0.0.1:11434
    roles:
      - chat
      - edit
      - apply
    defaultCompletionOptions:
      contextLength: 4096
      maxTokens: 512

Open your local configuration through Continue's configuration UI and adapt this block there. The current YAML reference requires name, version and schema and documents model roles and completion options. The Ollama provider reference documents provider: ollama and apiBase. The older config.json format is deprecated.

Select this local model in the editor before the first request. The context and output values are deliberate starting settings, not performance guarantees. If the runtime reports insufficient memory, reduce the context or choose another available model. Do not add a tool_use capability just to hide an unsupported-agent warning: configuration cannot create model behavior the artifact does not implement.

Provide the files and acceptance conditions the task needs

Scroll horizontally to see every column.

Context itemWhat to includeWhy
Target functionIts implementation and input/output contractGive the model the behavior actually being changed.
Relevant callerOne representative call siteExpose assumptions about types and error handling.
Existing testThe failing case and nearby test styleMake the requested outcome verifiable.
Project instructionsApplicable conventions and supported runtimeAvoid incompatible APIs or unwanted dependencies.
Change scopeAllowed files and explicit exclusionsKeep review focused on one approved task.

Attach only the context needed for the edit through the editor's available file-selection controls. Inspect the selected provider before sharing repository contents and omit credentials or unnecessary production values. Your direct local model path does not establish that unrelated telemetry, synchronization, remote MCP tools or documentation retrieval also stays local.

Use stable acceptance conditions instead of asking the model to improve code generally. For example: require a retry count to be an integer from zero through three, reject numeric strings and booleans, and preserve existing valid callers. Ask for the smallest change and an explanation of the behavior affected. For larger context, follow context-window planning.

Check behavior with tests that can reject a plausible patch

retry-contract.mjs: a runnable synthetic behavior fixture, Node standard library only
import assert from "node:assert/strict"

// Reference behavior for a small synthetic task, not a model-generated patch.
function validateRetryCount(value) {
  if (!Number.isInteger(value) || value < 0 || value > 3) {
    throw new RangeError("Retry count must be an integer from 0 to 3.")
  }
  return value
}

assert.equal(validateRetryCount(0), 0)
assert.equal(validateRetryCount(3), 3)
for (const invalid of [-1, 4, 1.5, "2", true, null]) {
  assert.throws(() => validateRetryCount(invalid), RangeError)
}
console.log("Synthetic retry-count contract verified.")

Run node retry-contract.mjs locally to verify the illustrated contract. This self-contained reference is not a test of an AI-generated repository patch. In the actual project, make its corresponding tests import the function the patch changes, including accepted boundaries and rejected types. Then run the repository's documented validation commands and inspect the diff.

The fixture uses Node's strict assertions to check both values and thrown errors. A model's explanation that tests pass is not execution evidence. Keep the actual command outcome, examine failures and review whether the patch changed unrelated behavior, added a dependency or weakened the checks just to satisfy the test.

Improve model choice without expanding permissions blindly

When an edit fails, give the assistant the specific failing output and relevant code, then request one bounded revision. Keep the prompt and expected behavior fixed when comparing models. Record answer usefulness, review effort and latency, using the benchmark method when timing matters. Faster text generation is not the same as producing a correct, reviewable change.

Chat, autocomplete and agent tool use are different tasks. Evaluate each before enabling it, and approve repository writes or commands according to your editor's controls and team policy. Tool-calling design explains the application-owned authorization loop. Review the complete data path if you later move from a direct local provider to a shared inference endpoint.

Keep shared inference and bounded CPU tools distinct from the editor

gpuOS can serve deployed open models from your connected GPU through its hosted gateway, which receives prompts and outputs. That route has a different data boundary from this direct loopback configuration. Use the OpenAI-compatible API guide when evaluating a shared application endpoint, and verify the client and model capabilities rather than assuming every editor role transfers unchanged.

cpuOS is complementary for an application-owned, approved calculation or data transformation. Its pilot runs trusted standard-library Python and Node jobs on a Docker worker, with bounded resources and no job network, repository checkout, uploaded files or package installation. It is not the runtime for this complete editor workflow. The cpuOS Python jobs tutorial shows the separate submit, poll and checked-result contract.

Questions

Do I need an OpenAI API key for this Continue and Ollama setup?
No. This example selects Continue's Ollama provider at a direct local endpoint with an already downloaded model. Check that the editor actually selects that provider and has no unintended cloud fallback.
Does assigning a model the chat or edit role enable autonomous tools?
No. Roles select editor functions; tool use requires a model and runtime that actually support it. Do not declare tool_use simply to bypass an unsupported capability warning.
Can cpuOS run my whole coding repository for this assistant?
No. Its current pilot provides bounded trusted Python and Node standard-library jobs, without repository checkout, networking, uploaded files or package installation. Run repository validation in your own approved development environment.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.