Define which assistant actions must work without internet
An offline AI assistant needs an installed runtime, local model files and a workflow whose required tools do not depend on external services. Start with one explicit action, such as summarizing a note already on your machine. Web search, a cloud-hosted model and a remote documentation tool change that action's data path even if the chat interface itself is a desktop app.
This guide uses a direct local Ollama server and a synthetic note to verify the path. A successful disconnected request establishes that tested path's availability, not a general guarantee about every extension or future configuration. Keep local AI architecture and local versus cloud processing separate from claims based only on where the interface is installed.
Prepare the runtime, model and optional interface while connected
Scroll horizontally to see every column.
| Download or record | Why it matters offline |
|---|---|
| Runtime installer and version | The model file alone does not include the inference application. |
| Exact local model tag and artifact | A catalog entry or configuration block does not download weights. |
| Editor extension or local UI | Installation and updates may need external access. |
| Any local embedding model | Document retrieval can invoke a separate model from chat. |
| Required notes and documentation | A remote URL is not an offline knowledge source. |
| Working settings and acceptance prompts | Reproduce the setup after a process or machine restart. |
ollama --version
ollama pull qwen3:0.6b
ollama listThe verified example model tag provides a small functional candidate. Confirm a synthetic request works before disconnecting, then close the assistant and test again after restart. That restart checks the installed files rather than only a model still resident in memory. Adapt the model to your task after the basic path works.
Disable cloud inference and tools in the actual server process
Ollama documents a local-only mode through disable_ollama_cloud in its server configuration or the OLLAMA_NO_CLOUD=1 environment variable. The cloud-disable FAQ says to restart the server and check its log confirmation. Apply a setting to the process that actually serves requests; setting an environment variable in a different shell does not change an already running desktop app or service.
{
"disable_ollama_cloud": true
}Preserve other properties when editing an existing server configuration. After restarting, check for Ollama cloud disabled: true in the server logs. Select the downloaded local model explicitly. Remove cloud fallback providers, web search and remote MCP servers from the assistant configuration, and inspect helper processes started by local tools. A tool launched on your machine can still make an external request.
When using Continue as an interface, its offline setup guide describes installing the extension in advance, selecting local models, disabling anonymous telemetry and restarting the editor. Review the editor's own telemetry, sync, updates and other extensions separately. That assistant setting does not configure every application on the machine.
Verify server placement and prevent accidental proxy routing
Keep the inference server bound to loopback for this same-machine test and connect to 127.0.0.1:11434. Verify that the expected local process owns the listening port and that its configured model is loaded there. Inspect outgoing connections from the server, interface and helper processes while a request runs. A localhost address identifies the first hop; a local proxy or router can forward work elsewhere.
Disable remote-device routing if your chosen interface provides it. Check HTTP proxy settings and inherited provider configurations too. The request below uses curl's --noproxy '*' to bypass configured proxies for this test. The curl manual documents that option. Bypassing a proxy for one request does not disable other applications' networking.
Use your platform's network controls or firewall policy to block external networking while retaining loopback. Then restart the runtime and interface with the same configuration. Save the settings and connection observations alongside the test outcome; successfully generating text does not itself prove that every process avoided external traffic.
Repeat a self-contained request with external networking blocked
# Run after the model is downloaded and external networking is blocked.
# --noproxy keeps this loopback request out of configured HTTP proxies.
curl --noproxy '*' --fail-with-body http://127.0.0.1:11434/api/chat \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3:0.6b",
"stream": false,
"think": false,
"messages": [{
"role": "user",
"content": "Use only this note: The workshop meets on Tuesday at 14:00 in Room 2. When and where does it meet? If the note does not answer, say unknown."
}],
"options": {"num_ctx": 2048, "num_predict": 128, "temperature": 0}
}' > offline-response.jsonInspect offline-response.json: require a completed response, no API error and an answer that preserves Tuesday, 14:00 and Room 2. The Ollama chat reference defines the returned message and completion fields. Also ask a question the note does not answer, such as who booked the room, and check that the assistant admits the missing information.
Test each required feature separately: first ordinary chat, then reading a permitted local note, then local document retrieval if configured. A successful chat request does not establish that an embedding model, file reader or citation tool is available offline. Keep the model's answer-quality check separate from the connection check, and record any attempted external requests even when the firewall blocked them.
Keep the offline workflow reproducible after changes
- Save the runtime and model versions, server settings, interface provider and the exact tested features.
- Repeat the disconnected test after changing a model, extension, retrieval backend or remote tool configuration.
- Check that a missing document yields an explicit failure rather than a silent web fallback.
- Keep source passages beside summaries and inspect factual details before acting on them.
- Plan controlled connected periods for updates and downloads, then recheck the intended offline configuration.
gpuOS's current service uses a hosted gateway that receives prompts and outputs before routing inference to your connected GPU. It is a different route from this directly local offline assistant. cpuOS likewise uses a hosted control plane for bounded trusted-code jobs. Neither service is an offline backend merely because you own the worker hardware. For an editor-specific local workflow, follow local AI coding assistance.