gpuos

Use cases · 5 min read · updated Oct 7, 2026

Build an offline AI assistant and verify its network path

Prepare a local AI assistant for offline use. Download models, disable cloud features and remote tools, then verify the complete workflow without internet.

On this page

Define which assistant actions must work without internet

An offline AI assistant needs an installed runtime, local model files and a workflow whose required tools do not depend on external services. Start with one explicit action, such as summarizing a note already on your machine. Web search, a cloud-hosted model and a remote documentation tool change that action's data path even if the chat interface itself is a desktop app.

This guide uses a direct local Ollama server and a synthetic note to verify the path. A successful disconnected request establishes that tested path's availability, not a general guarantee about every extension or future configuration. Keep local AI architecture and local versus cloud processing separate from claims based only on where the interface is installed.

Prepare the runtime, model and optional interface while connected

Scroll horizontally to see every column.

Download or recordWhy it matters offline
Runtime installer and versionThe model file alone does not include the inference application.
Exact local model tag and artifactA catalog entry or configuration block does not download weights.
Editor extension or local UIInstallation and updates may need external access.
Any local embedding modelDocument retrieval can invoke a separate model from chat.
Required notes and documentationA remote URL is not an offline knowledge source.
Working settings and acceptance promptsReproduce the setup after a process or machine restart.
Preparation before disconnecting external networking
ollama --version
ollama pull qwen3:0.6b
ollama list

The verified example model tag provides a small functional candidate. Confirm a synthetic request works before disconnecting, then close the assistant and test again after restart. That restart checks the installed files rather than only a model still resident in memory. Adapt the model to your task after the basic path works.

Disable cloud inference and tools in the actual server process

Ollama documents a local-only mode through disable_ollama_cloud in its server configuration or the OLLAMA_NO_CLOUD=1 environment variable. The cloud-disable FAQ says to restart the server and check its log confirmation. Apply a setting to the process that actually serves requests; setting an environment variable in a different shell does not change an already running desktop app or service.

Merge this property into the existing ~/.ollama/server.json, then restart Ollama
{
  "disable_ollama_cloud": true
}

Preserve other properties when editing an existing server configuration. After restarting, check for Ollama cloud disabled: true in the server logs. Select the downloaded local model explicitly. Remove cloud fallback providers, web search and remote MCP servers from the assistant configuration, and inspect helper processes started by local tools. A tool launched on your machine can still make an external request.

When using Continue as an interface, its offline setup guide describes installing the extension in advance, selecting local models, disabling anonymous telemetry and restarting the editor. Review the editor's own telemetry, sync, updates and other extensions separately. That assistant setting does not configure every application on the machine.

Verify server placement and prevent accidental proxy routing

Keep the inference server bound to loopback for this same-machine test and connect to 127.0.0.1:11434. Verify that the expected local process owns the listening port and that its configured model is loaded there. Inspect outgoing connections from the server, interface and helper processes while a request runs. A localhost address identifies the first hop; a local proxy or router can forward work elsewhere.

Disable remote-device routing if your chosen interface provides it. Check HTTP proxy settings and inherited provider configurations too. The request below uses curl's --noproxy '*' to bypass configured proxies for this test. The curl manual documents that option. Bypassing a proxy for one request does not disable other applications' networking.

Use your platform's network controls or firewall policy to block external networking while retaining loopback. Then restart the runtime and interface with the same configuration. Save the settings and connection observations alongside the test outcome; successfully generating text does not itself prove that every process avoided external traffic.

Repeat a self-contained request with external networking blocked

A loopback-only client request with the complete synthetic note inline
# Run after the model is downloaded and external networking is blocked.
# --noproxy keeps this loopback request out of configured HTTP proxies.
curl --noproxy '*' --fail-with-body http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3:0.6b",
    "stream": false,
    "think": false,
    "messages": [{
      "role": "user",
      "content": "Use only this note: The workshop meets on Tuesday at 14:00 in Room 2. When and where does it meet? If the note does not answer, say unknown."
    }],
    "options": {"num_ctx": 2048, "num_predict": 128, "temperature": 0}
  }' > offline-response.json

Inspect offline-response.json: require a completed response, no API error and an answer that preserves Tuesday, 14:00 and Room 2. The Ollama chat reference defines the returned message and completion fields. Also ask a question the note does not answer, such as who booked the room, and check that the assistant admits the missing information.

Test each required feature separately: first ordinary chat, then reading a permitted local note, then local document retrieval if configured. A successful chat request does not establish that an embedding model, file reader or citation tool is available offline. Keep the model's answer-quality check separate from the connection check, and record any attempted external requests even when the firewall blocked them.

Keep the offline workflow reproducible after changes

  • Save the runtime and model versions, server settings, interface provider and the exact tested features.
  • Repeat the disconnected test after changing a model, extension, retrieval backend or remote tool configuration.
  • Check that a missing document yields an explicit failure rather than a silent web fallback.
  • Keep source passages beside summaries and inspect factual details before acting on them.
  • Plan controlled connected periods for updates and downloads, then recheck the intended offline configuration.

gpuOS's current service uses a hosted gateway that receives prompts and outputs before routing inference to your connected GPU. It is a different route from this directly local offline assistant. cpuOS likewise uses a hosted control plane for bounded trusted-code jobs. Neither service is an offline backend merely because you own the worker hardware. For an editor-specific local workflow, follow local AI coding assistance.

Questions

Can I install and use a local assistant after disconnecting internet?
Prepare the runtime, exact model files, interface and any retrieval models before disconnecting. Then restart and test the complete required workflow with external networking blocked and loopback available.
Does connecting to localhost prove that inference stayed on my machine?
No. A local proxy or routing service can forward a request. Check which process owns the endpoint, its provider configuration and outgoing connections while testing with external access blocked.
Is gpuOS an offline inference service because I own the GPU?
No. gpuOS uses a hosted gateway that receives prompts and outputs. This guide's offline workflow uses a separate direct local runtime and locally available tools.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.