Separate the inference location from the application's access layer
Local AI runs inference on a machine you use or manage. Cloud AI sends requests to a remotely provided inference service. A private server, a rented GPU and a hosted gateway can combine aspects of both. Describe the machine executing the model and the services carrying its requests instead of relying only on ‘local’ or ‘cloud’ as labels.
Ollama demonstrates that this is a deployment choice rather than a fixed property of one application: it supports local models and cloud models. Its quickstart describes selecting each route. Keep the route explicit in your configuration so an application does not silently change its processing location when a model name changes.
Start from the task, traffic pattern and quality requirement
Scroll horizontally to see every column.
| Workload | What to evaluate locally | What to evaluate remotely |
|---|---|---|
| Personal drafting | Useful quality and acceptable wait on existing hardware | Quality improvement worth adding a remote dependency |
| Internal document questions | Retrieval quality, context capacity and shared access | Allowed data path and evidence handling |
| Scheduled extraction | Completion time within the batch window | Throughput limits and predictable request cost |
| Interactive coding | Tool behavior, repository context and time to first output | Client compatibility and complete tool-round latency |
| Unpredictable traffic | Capacity and queuing at the busiest period | Quota, service limits and failure behavior |
Write an acceptance test before choosing a deployment. For extraction, check required fields and values. For document questions, check whether citations support the answer. For coding assistance, run the relevant checks on proposed edits. If a model fails the useful task, cheaper requests or higher token throughput do not fix that failure.
Compare total workload cost, including idle capacity and retries
For local inference, list hardware purchase or rental, energy, storage, maintenance and the time someone spends operating it. If you already own the machine, evaluate both the extra running cost and what other work it could perform. A GPU that serves one short request a day and one kept busy by batch jobs have different economics.
For a cloud service, use its current charging method and your measured input/output volume. Include repeated requests, long prompts, failed operations that are billed and any separate tool or storage charges that apply. This guide does not assume token pricing, a free tier or a particular provider's retention policy.
Compare the cost of completing the same accepted tasks over a stated period. Record request volume and concurrency alongside that result. A break-even estimate is useful only with its assumptions visible: changing the model, quality threshold or traffic pattern can change the conclusion. Revisit the calculation after observing actual usage.
Map every component that receives prompts or documents
Trace the interface, inference endpoint, retrieval index, tools, logs and saved outputs. In a local-only setup, keep those components on the intended device or network and test their operation together. In a remote setup, identify the services receiving each part of the request and review their applicable processing and retention terms.
A desktop app can offer local and remote modes. LM Studio's privacy policy explicitly distinguishes local-model use from optional cloud services. Review the selected features and connected tools rather than generalizing one mode's data flow to all uses of the application.
gpuOS runs Ollama on your connected GPU machine and routes requests through its hosted gateway, which receives prompts and outputs. It provides hosted access to inference hardware you control. If your requirement is that the entire request stays offline on one device, evaluate a direct local workflow instead. The local AI guide shows how to establish that starting point.
Decide who owns availability, capacity and application recovery
With a machine you operate, someone must handle service startup, model downloads, resource contention and hardware failure. A local desktop that sleeps between sessions can work for personal use while being unsuitable as an always-available team dependency. Test the actual deployment schedule and restart behavior.
A remote service moves some operational responsibility to its provider but still leaves application work: limits, authentication, request timeouts, retries and error handling. Specify how your application behaves when the endpoint is unavailable. If you keep a fallback model, evaluate its quality and make its processing location explicit to the application and its users.
cpuOS addresses a different part of an application: trusted Python and Node.js standard-library jobs executed in Docker on connected CPU machines. It is not an LLM inference backend. Keep deterministic preprocessing or report generation separate from the decision about where the language model runs.
Run a small comparison before buying capacity or migrating traffic
- Select representative tasks, input sizes and accepted outputs, including difficult cases.
- Run each candidate with the same task requirements, while documenting model and configuration differences.
- Measure complete task latency, failures and cost inputs, including retries and tool calls.
- Test expected concurrency and a period when the primary route is unavailable.
- Choose the route that meets the quality, data-flow and service requirements, and save the evidence for the next review.
Use the benchmark guide to label timing boundaries and the private RAG guide for document workflows. A mixed deployment can be useful when each task has an explicit route and fallback rule. Evaluate that routing policy as part of the application, rather than assuming one location is best for every request.