gpuos

Getting started · 4 min read · updated Oct 7, 2026

Local AI: what it means and how to run your first model

Get started with local AI using Ollama or LM Studio. Run a model, test a useful task, check memory and trace where your prompts and files go.

On this page

Local AI describes where inference runs

Local AI means running an AI model on hardware you use or manage, such as a laptop, workstation or private server. For a local language model, the software loads model weights and generates answers on that machine. This guide uses ‘local AI’ as a general term; LocalAI is also the name of a separate software project.

A local inference engine is one component of an application. Its chat interface, document index, extensions and remote tools can each have a different data path. Decide whether you need on-device inference, a shared server you control or an entirely offline workflow. Those requirements lead to different setups, even when the same model can serve all three.

Start with one task and a model your machine can load

Choose a small task with an answer you can check: rewrite a paragraph, extract three fields from synthetic notes or explain a short function. Begin with a short prompt so you can separate installation problems from input-size problems. Save that prompt and its expected result before adding documents or a coding-agent integration.

  • Check free disk space for the download and available RAM or GPU memory for inference.
  • Choose an available model variant suited to your language and task. A smaller model is a starting point to evaluate, not a promise of identical quality.
  • Record the exact model name and precision so you can repeat the experiment later.
  • Keep the first input free of real customer data; introduce your actual workflow after checking its complete request path.

The VRAM guide and quantization guide help estimate GPU capacity. If your machine has no suitable accelerator, use the local LLM without a GPU guide to evaluate CPU inference and realistic task sizes.

Download a model and make a first request with Ollama

Install and start Ollama using its official quickstart. Select a local model for this exercise. The commands below assume Ollama is running and your machine has capacity for Qwen3 8B. The download needs connectivity; the test prompt is synthetic. Choose a smaller supported model if this variant does not fit.

Download, run and inspect a local model
ollama pull qwen3:8b
ollama run qwen3:8b "Summarize in one sentence: The demo team moved its planning meeting from Monday to Wednesday."
ollama ps

Check that the answer preserves Wednesday and does not add a reason for the change. Then try a question whose answer is missing from the text and inspect whether the model acknowledges that. These two checks are more informative than a greeting: they exercise both following evidence and handling missing information.

The Ollama CLI reference documents downloading, running and inspecting models. If a request fails, first check whether the download finished and the service is running. If it succeeds slowly, record the loaded processor placement and test a shorter task before changing several settings at once.

Choose a chat interface or connect your own application

Scroll horizontally to see every column.

Your next stepA practical workflowCheck before expanding it
Personal chatUse a local chat interface with a downloaded modelSelected model and where chat history is saved
Document questionsAdd relevant passages with source labelsDocument processing and any embedding service
Application integrationSend a short request to the engine's local APIEndpoint, model id and response handling
Team accessServe the model through an authenticated shared routeData flow, permissions and concurrent capacity

For a graphical starting point, LM Studio offers a desktop chat interface; its offline operation documentation explains which activities need connectivity. Compare the workflows in Ollama vs LM Studio rather than switching tools merely because a different interface looks easier.

A successful local chat does not establish API compatibility with every client. Test the route your application uses, then add history, streaming or tools one feature at a time. Keep a simple working request available for diagnosing integration errors.

Trace the complete path of prompts, files and answers

Write down the route from the interface to the model and back. Include external search, tool servers, embedding endpoints, analytics, saved histories and backups when they are part of your application. A local model does not determine where those other components run. For an offline requirement, test the completed workflow without connectivity, including follow-up questions.

Ollama also offers cloud models, so verify the selected model and route rather than treating the software name as proof of local inference. In gpuOS, Ollama runs on your connected GPU machine, while prompts and outputs pass through the hosted gpuOS gateway. That makes gpuOS a shared hosted access layer over your hardware, not an offline-only application.

Keep a short evaluation before relying on the model

Collect ordinary inputs, missing information and a few difficult examples from your intended task. Record whether each answer is correct, whether it follows the requested format and how long it takes. Re-run the same set after changing the model, prompt or precision. For a repeatable timing method, use the local LLM benchmark guide.

When one user's workflow works, decide what changes for shared use: access keys, simultaneous requests, capacity and support. The self-hosted API guide covers the gpuOS route. Keep the local test as a baseline so you can distinguish model behavior from the extra application and network layers.

Questions

Does local AI mean the whole application works offline?
It describes where model inference runs. Downloads, cloud models, search, tool servers or other application services can still need connectivity. Test the complete workflow offline if that is a requirement.
Do I need an NVIDIA GPU to try local AI?
Not for every runtime or model. Supported hardware varies, and CPU inference is possible in suitable setups. Choose a model your machine can load, then evaluate response quality and latency for your task.
Is gpuOS an entirely local AI chat application?
gpuOS runs Ollama inference on your connected GPU machine and provides a hosted API gateway for shared access. Prompts and outputs pass through that gateway, so its complete request path is not offline or confined to one device.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.