30-SECOND SUMMARY
What to take away
- Local execution gives you more control over data paths, but it is not automatically secure.
- Start with one runtime and one small 3B–8B quantized model.
- Evaluate models with your own documents and repeatable questions instead of relying only on leaderboards.
From prompt to local answer
External services are optional, not part of the default path.
- 01INPUT
Question or document
- 02RUNTIME
Ollama or LM Studio
- 03MODEL
On-device inference
- 04OUTPUT
Human verification
What can my computer run?
We use the lowest of the three components as a conservative estimate.
- Model candidates
- 3B–8B Q4
- Suggested work
- Chat, translation, and summaries
- Limiting components
- CPU · RAM · GPU
This is a planning estimate, not a compatibility guarantee.
| Tier | CPU example | RAM | GPU example | Model candidates |
|---|---|---|---|---|
| Experiment | Older; verify support | 8GB | CPU execution | 1B–3B Q4 |
| Starter | Recent 4–6 cores; M1/M2 | 16GB | Integrated; 4–6GB VRAM | 3B–8B Q4 |
| Working | Recent 8+ cores; Pro class | 32GB | 8–12GB VRAM | 8B–14B Q4 |
| Expanded | High-end 12+ cores; Max class | 64GB+ | 16–24GB+ VRAM | 14B–32B Q4 |
Important: This is not a universal minimum or purchase guarantee. CPU generation, instruction support, bandwidth, context, quantization, and runtime all matter. Test a small model before buying hardware.
What local AI actually means
Cloud AI sends a request to infrastructure operated by a service provider. Local AI stores the model weights and runs inference on your device. Once the runtime and model are downloaded, basic chat and document work can operate without an internet connection.
The trade-off is responsibility. You manage model files, storage, memory, updates, logs, network exposure, and licensing yourself. The word local describes inference location; it does not prove that every plugin, backup, or connected feature stays offline.
Before treating a workflow as local, list its optional network paths: model downloads, a `:cloud` tag, web search, editor extensions, sync folders, and telemetry. Run one harmless prompt with the network disconnected and record what still works. This verifies the inference path, but not every copy that may remain in logs or backups.
When local is a good fit
Local models are useful for private drafts, offline work, repetitive development tests, and workflows where you want to control exactly which model and prompt are used. They are especially attractive when a person can compare the output with a known source.
Cloud services remain stronger when you need frontier reasoning, very large context windows, managed web search, or complex tool use. A hybrid workflow can keep drafts local and reserve an approved service for current research or a final review.
Choose the boundary from the consequence of an error. A private draft or repeatable development test can be a strong local candidate; a task needing current web evidence, managed collaboration, or a high-stakes answer needs controls beyond where the model runs. Keep a human source-check step for facts and decisions.
| Factor | Local AI | Cloud AI |
|---|---|---|
| Data path | Can stay on-device | Sent to provider infrastructure |
| Internet | Optional after download | Usually required |
| Cost | Hardware, power, maintenance | Subscription or usage |
| Management | You maintain it | Provider maintains infrastructure |
Choose one runtime
Ollama is a straightforward choice for terminal and API workflows. LM Studio is easier when you prefer a desktop interface for finding and loading models. llama.cpp offers deeper control over backends and performance settings.
Do not install everything at once. Fix the runtime first, then compare models so that you can isolate the cause of any change. Complete load, chat, stop, restart, and removal with one model before adding another runtime.
Use one acceptance test before comparing runtimes: load the same model, run the same five prompts, note the model path, context setting, first-response wait, and where logs live. Change either runtime or model—not both—so a slower or less accurate result has an explainable cause.
Start with a small model
Model labels such as 3B, 8B, and 14B describe approximate parameter scale. Larger is not always better for your task and increases memory pressure. Treat 3B–8B as a possible starting range, not a compatibility promise.
Begin with the smallest instruct model that can complete the task. Leave memory for the operating system, runtime, and context cache; fitting the file on disk is not enough. Check the model card, license, exact tag, and download size before installation.
Watch the process during a short and a long prompt. If memory pressure or CPU/GPU offloading changes sharply, reduce model size or context before assuming that a larger download is an improvement. Adopt a larger candidate only when the same known-answer test shows a useful quality gain.
Run the first conversation
Install one runtime from its official source, open a new terminal, and run a currently listed local model. A command such as `ollama run gemma3:4b` is an example, not a permanent default; verify the exact tag in the model library before downloading.
Use a short public question whose answer you can check. A completed response and a second prompt show that the basic chat loop works, while `ollama ls` confirms the model is stored and `ollama ps` reports a model that is currently loaded.
If the command fails, do not begin with a reinstall. Check free storage, whether the local service is running, and whether the model tag is exact. Preserve the error message and change one condition at a time so the eventual fix can be repeated.
ollama run gemma3:4b
>>> Explain local AI in three sentences.
/byeRun a repeatable first test
Prepare five questions whose answers you already know. Include factual extraction, summarization, formatting, an answer-not-found case, and a repeated prompt. Keep the source, prompt, model, context, and generation settings fixed.
Record the runtime version, full model name, quantization, context setting, time to first output, and whether the answer remained consistent across three runs. Score factual match, instruction following, format validity, and unsupported claims separately.
Write a pass rule before starting: for example, all five answers match the supplied evidence, the required format parses, and three runs do not change the material conclusion. A fluent answer that invents a detail is a failed test; if every candidate fails, narrow the task instead of selecting the least bad output.
- Use public or synthetic data first
- Keep the prompt and settings fixed
- Check facts against the source
- Record memory use and failures
- Test offline behavior before adding private documents
Close the first week with evidence
The first week is for a repeatable baseline, not the largest model your computer can launch. Keep the runtime version, exact model tag, context setting, five test prompts, observed memory use, and completion times in one record.
Do not introduce private material until you have checked logs, network behavior, backup folders, connected tools, and model licensing with public data. Test stopping and deleting a model so that the cleanup path is understood before the stakes rise.
At the end of the week, ask whether one real task was completed accurately and consistently at an acceptable wait time. Repeat the baseline after updates, and expand scope only when the same checks still pass. This turns a successful demo into an operating decision you can defend.
Frequently asked questions
Can local AI work completely offline?
Yes for basic inference after the runtime and model are installed. Downloads, updates, cloud models, and web search still require a network. For a practical check, follow the “What local AI actually means” section, change one condition at a time, and record the result.
Does a free model allow commercial use?
Not necessarily. Check the model card, license, and your organization’s policy. For a practical check, follow the “When local is a good fit” section, change one condition at a time, and record the result.
Do I need a dedicated GPU?
No for a small CPU-based test, but generation may be slow. Supported GPU acceleration can materially improve the experience. For a practical check, follow the “Choose one runtime” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Quickstart Ollama Model Library LM Studio Documentation llama.cpp Repository