PRIVATE · PRACTICAL · LOCAL
Run AI locally,
on your computer.
Field notes from setting up local AI on existing PCs: the questions, failures, fixes, and repeatable checks—not only the commands that worked.
$ ollama run gemma3:4b
pulling manifest ··· done
loading model locally ··· done
› Where is my document sent?
This conversation is processed on this computer.
SERIES 01
First field note
OPSFIELD NOTE 01
Why We Connected Several PCs Instead of Buying One Bigger Machine
The project began when chat, document processing, and background jobs competed on one machine.
Read field note →What can my computer run?
We use the lowest of the three components as a conservative estimate.
Chat, translation, and summaries
Limiting components: CPU · RAM · GPU
A planning estimate; actual speed and compatibility depend on the model and runtime.
See the full tier guide →FIELD NOTES & GUIDES
Follow the setup journey
From the first install to operating several PCs, with related reference guides alongside the field notes.
Setup journey
The questions, failures, fixes, and checks in the order they happened.
Why We Used SQLite for a Local AI Job Coordinator
Preserve jobs, workers, and leases while changing ownership atomically.
What a Worker Must Report Beyond ‘Online’
A useful heartbeat includes roles, ready models, RAM, thermal state, power, user activity, and current work.
Why We Connected Several PCs Instead of Buying One Bigger Machine
The project began when chat, document processing, and background jobs competed on one machine.
Classify Existing Hardware by Role, Not Product Tier
Record memory headroom, cooling, power, network stability, and user conflicts before assigning work.
Related guides
Supporting material for understanding or reproducing the field notes.
Used PC and GPU Buying Checklist for Local AI
Verify VRAM, compatibility, power, cooling, condition, and a real model run before buying used hardware.
Secure a Local AI API
Keep localhost as the default and add authentication, TLS, network policy, limits, and monitoring before remote access.
Connect Open WebUI to Ollama
Run Open WebUI with Docker, connect it to local Ollama, and understand ports, accounts, and persistent data.
Build Your First Local AI App with the Ollama API
Call localhost, manage chat messages, choose streaming, request structured JSON, and handle failures safely.
Run Local AI on an Old Laptop
Check CPU features, RAM, storage, heat, and a small model before replacing an older computer.
Mac vs NVIDIA PC for Local AI
Compare unified memory, dedicated VRAM, CUDA and Metal ecosystems, mobility, upgrades, and real workload fit.
Local Speech-to-Text with Whisper
Install ffmpeg and Whisper, transcribe a short public recording, and review names, numbers, and sensitive artifacts.
Local RAG: Search Your Documents with a Local LLM
Understand extraction, chunking, embeddings, retrieval, grounded answers, evaluation, and deletion in a local RAG pipeline.
Benchmark Local LLM Speed Correctly
Separate loading, time to first token, and generation throughput under repeatable conditions.
How to Evaluate a Local LLM for Your Real Work
Build a fixed prompt set, repeat tests, score quality and operations separately, and define a rational model-switching threshold.
Summarize PDFs with Local AI
Extract clean text, preserve page evidence, summarize in stages, verify claims, and delete every derived copy.
LM Studio Setup: Run Your First Local LLM
Check requirements, find a GGUF model, choose quantization, load it into memory, and verify offline operation.
Install llama.cpp and Run a GGUF Model
Choose an official installation path, run a licensed GGUF in the CLI, then start a local compatible server.
How to Compare Local LLMs for Korean
Evaluate Korean comprehension, summaries, register, evidence, and consistency with a fixed task set.
GGUF Quantization: Q4 vs Q5 for Local LLMs
Learn what GGUF stores, how quantization trades memory for precision, and how to compare Q4 and Q5 on your hardware.
Context Length and KV Cache Explained
Understand how long context affects memory and latency, then set only what the task needs.
Install Ollama and Run Your First Local Model
Install Ollama on macOS, Windows, or Linux, run a model, verify GPU loading, and troubleshoot common failures.
How Much RAM and VRAM Do Local LLMs Need?
Estimate model-weight memory, understand context overhead, and choose a realistic model size for your hardware.
Local AI Privacy and Security Checklist
Map document paths, logs, backups, cloud features, model files, and network exposure before using sensitive data.
Local AI for Beginners: Where Should You Start?
A practical introduction to local AI, from choosing a runtime and model to testing quality on your own computer.
Category images: Unsplash
READING ORDER
Follow the setup
in sequence.
- 01
Inventory the hardware
Record memory, power, cooling, networking, and when each PC is available.
- 02
Run the first model
Establish a repeatable baseline with one runtime and one small model.
- 03
Connect the devices
Separate localhost and network problems, then restrict access to approved machines.
- 04
Assign the work
Split interactive and background roles, priorities, and ownership.
- 05
Drill recovery
Create reboots and disconnects deliberately, then verify requests and results survive.
“
We document what is reproducible, not merely fast.
Every guide includes its verification scope, update date, and primary sources.