Start with a reviewed edit in an ordinary repository
A local AI coding assistant combines an editor, repository context and a locally served model. Start with a small edit whose behavior you can verify: explain a function, add an input check or repair one failing test. Keep your existing version-control and review workflow. A model response is a proposed change; your project's tests and human review establish whether that change belongs in the codebase.
This setup uses Continue with a direct local Ollama endpoint. It demonstrates chat and editing, without assuming that the chosen model supports autonomous tool use. The local AI guide explains the component choices. If strict offline operation matters, also complete the offline assistant verification for the editor, server and any tools.
Install the interface and pull the exact coding model
ollama --version
ollama pull qwen2.5-coder:7b
ollama listInstall Continue in your editor and use the runtime's supported installation procedure. The Ollama model listing verifies this tag, while Continue's Ollama guide explains that a model configuration does not download its weights. Check the exact installed name rather than mixing a catalog ID, display name and runtime tag.
This model is an example, not an assertion that it is the best available coding model or fits every laptop. Choose a smaller supported variant when necessary and measure your intended prompts. Coding requests include file excerpts, instructions and an output allowance, so memory and latency should be checked with that actual context rather than a short greeting.
Use the current YAML provider configuration
name: Local coding review
version: 1.0.0
schema: v1
models:
- name: Local Qwen coder
provider: ollama
model: qwen2.5-coder:7b
apiBase: http://127.0.0.1:11434
roles:
- chat
- edit
- apply
defaultCompletionOptions:
contextLength: 4096
maxTokens: 512Open your local configuration through Continue's configuration UI and adapt this block there. The current YAML reference requires name, version and schema and documents model roles and completion options. The Ollama provider reference documents provider: ollama and apiBase. The older config.json format is deprecated.
Select this local model in the editor before the first request. The context and output values are deliberate starting settings, not performance guarantees. If the runtime reports insufficient memory, reduce the context or choose another available model. Do not add a tool_use capability just to hide an unsupported-agent warning: configuration cannot create model behavior the artifact does not implement.
Provide the files and acceptance conditions the task needs
Scroll horizontally to see every column.
| Context item | What to include | Why |
|---|---|---|
| Target function | Its implementation and input/output contract | Give the model the behavior actually being changed. |
| Relevant caller | One representative call site | Expose assumptions about types and error handling. |
| Existing test | The failing case and nearby test style | Make the requested outcome verifiable. |
| Project instructions | Applicable conventions and supported runtime | Avoid incompatible APIs or unwanted dependencies. |
| Change scope | Allowed files and explicit exclusions | Keep review focused on one approved task. |
Attach only the context needed for the edit through the editor's available file-selection controls. Inspect the selected provider before sharing repository contents and omit credentials or unnecessary production values. Your direct local model path does not establish that unrelated telemetry, synchronization, remote MCP tools or documentation retrieval also stays local.
Use stable acceptance conditions instead of asking the model to improve code generally. For example: require a retry count to be an integer from zero through three, reject numeric strings and booleans, and preserve existing valid callers. Ask for the smallest change and an explanation of the behavior affected. For larger context, follow context-window planning.
Check behavior with tests that can reject a plausible patch
import assert from "node:assert/strict"
// Reference behavior for a small synthetic task, not a model-generated patch.
function validateRetryCount(value) {
if (!Number.isInteger(value) || value < 0 || value > 3) {
throw new RangeError("Retry count must be an integer from 0 to 3.")
}
return value
}
assert.equal(validateRetryCount(0), 0)
assert.equal(validateRetryCount(3), 3)
for (const invalid of [-1, 4, 1.5, "2", true, null]) {
assert.throws(() => validateRetryCount(invalid), RangeError)
}
console.log("Synthetic retry-count contract verified.")Run node retry-contract.mjs locally to verify the illustrated contract. This self-contained reference is not a test of an AI-generated repository patch. In the actual project, make its corresponding tests import the function the patch changes, including accepted boundaries and rejected types. Then run the repository's documented validation commands and inspect the diff.
The fixture uses Node's strict assertions to check both values and thrown errors. A model's explanation that tests pass is not execution evidence. Keep the actual command outcome, examine failures and review whether the patch changed unrelated behavior, added a dependency or weakened the checks just to satisfy the test.
Improve model choice without expanding permissions blindly
When an edit fails, give the assistant the specific failing output and relevant code, then request one bounded revision. Keep the prompt and expected behavior fixed when comparing models. Record answer usefulness, review effort and latency, using the benchmark method when timing matters. Faster text generation is not the same as producing a correct, reviewable change.
Chat, autocomplete and agent tool use are different tasks. Evaluate each before enabling it, and approve repository writes or commands according to your editor's controls and team policy. Tool-calling design explains the application-owned authorization loop. Review the complete data path if you later move from a direct local provider to a shared inference endpoint.
Keep shared inference and bounded CPU tools distinct from the editor
gpuOS can serve deployed open models from your connected GPU through its hosted gateway, which receives prompts and outputs. That route has a different data boundary from this direct loopback configuration. Use the OpenAI-compatible API guide when evaluating a shared application endpoint, and verify the client and model capabilities rather than assuming every editor role transfers unchanged.
cpuOS is complementary for an application-owned, approved calculation or data transformation. Its pilot runs trusted standard-library Python and Node jobs on a Docker worker, with bounded resources and no job network, repository checkout, uploaded files or package installation. It is not the runtime for this complete editor workflow. The cpuOS Python jobs tutorial shows the separate submit, poll and checked-result contract.