PRIVATE · PRACTICAL · LOCAL

Run AI locally,
on your computer.

Field notes from setting up local AI on existing PCs: the questions, failures, fixes, and repeatable checks—not only the commands that worked.

local-ai — zsh

$ ollama run gemma3:4b

pulling manifest ··· done

loading model locally ··· done

› Where is my document sent?

This conversation is processed on this computer.

● NETWORK 0 KB
Run it ourselvesVerify in EnglishNo inflated numbersPublish update dates
HARDWARE CHECK 01

What can my computer run?

We use the lowest of the three components as a conservative estimate.

Estimated tierStarter3B–8B Q4

Chat, translation, and summaries
Limiting components: CPU · RAM · GPU

A planning estimate; actual speed and compatibility depend on the model and runtime.

See the full tier guide →

FIELD NOTES & GUIDES

Follow the setup journey

From the first install to operating several PCs, with related reference guides alongside the field notes.

Published
24 / 24 GUIDES
01 · FIELD NOTES

Setup journey

The questions, failures, fixes, and checks in the order they happened.

02 · REFERENCES

Related guides

Supporting material for understanding or reproducing the field notes.

HardwareIntermediate · Selection guide

Used PC and GPU Buying Checklist for Local AI

Verify VRAM, compatibility, power, cooling, condition, and a real model run before buying used hardware.

Use reference ↗
SecurityIntermediate · Security check

Secure a Local AI API

Keep localhost as the default and add authentication, TLS, network policy, limits, and monitoring before remote access.

Use reference ↗
ToolsBeginner · Step by step

Connect Open WebUI to Ollama

Run Open WebUI with Docker, connect it to local Ollama, and understand ports, accounts, and persistent data.

Use reference ↗
BuildIntermediate · Hands-on build

Build Your First Local AI App with the Ollama API

Call localhost, manage chat messages, choose streaming, request structured JSON, and handle failures safely.

Use reference ↗
StartBeginner · Step by step

Run Local AI on an Old Laptop

Check CPU features, RAM, storage, heat, and a small model before replacing an older computer.

Use reference ↗
HardwareIntermediate · Selection guide

Mac vs NVIDIA PC for Local AI

Compare unified memory, dedicated VRAM, CUDA and Metal ecosystems, mobility, upgrades, and real workload fit.

Use reference ↗
BuildBeginner · Hands-on build

Local Speech-to-Text with Whisper

Install ffmpeg and Whisper, transcribe a short public recording, and review names, numbers, and sensitive artifacts.

Use reference ↗
BuildIntermediate · Hands-on build

Local RAG: Search Your Documents with a Local LLM

Understand extraction, chunking, embeddings, retrieval, grounded answers, evaluation, and deletion in a local RAG pipeline.

Use reference ↗
EvaluateIntermediate · Compare & evaluate

Benchmark Local LLM Speed Correctly

Separate loading, time to first token, and generation throughput under repeatable conditions.

Use reference ↗
EvaluateIntermediate · Compare & evaluate

How to Evaluate a Local LLM for Your Real Work

Build a fixed prompt set, repeat tests, score quality and operations separately, and define a rational model-switching threshold.

Use reference ↗
BuildIntermediate · Hands-on build

Summarize PDFs with Local AI

Extract clean text, preserve page evidence, summarize in stages, verify claims, and delete every derived copy.

Use reference ↗
ToolsBeginner · Step by step

LM Studio Setup: Run Your First Local LLM

Check requirements, find a GGUF model, choose quantization, load it into memory, and verify offline operation.

Use reference ↗
ToolsBeginner · Step by step

Install llama.cpp and Run a GGUF Model

Choose an official installation path, run a licensed GGUF in the CLI, then start a local compatible server.

Use reference ↗
ModelsIntermediate · Compare & evaluate

How to Compare Local LLMs for Korean

Evaluate Korean comprehension, summaries, register, evidence, and consistency with a fixed task set.

Use reference ↗
ModelsIntermediate · Practical guide

GGUF Quantization: Q4 vs Q5 for Local LLMs

Learn what GGUF stores, how quantization trades memory for precision, and how to compare Q4 and Q5 on your hardware.

Use reference ↗
ModelsIntermediate · Practical guide

Context Length and KV Cache Explained

Understand how long context affects memory and latency, then set only what the task needs.

Use reference ↗
ToolsBeginner · Step by step

Install Ollama and Run Your First Local Model

Install Ollama on macOS, Windows, or Linux, run a model, verify GPU loading, and troubleshoot common failures.

Use reference ↗
HardwareIntermediate · Selection guide

How Much RAM and VRAM Do Local LLMs Need?

Estimate model-weight memory, understand context overhead, and choose a realistic model size for your hardware.

Use reference ↗
SecurityIntermediate · Security check

Local AI Privacy and Security Checklist

Map document paths, logs, backups, cloud features, model files, and network exposure before using sensitive data.

Use reference ↗
StartBeginner · Step by step

Local AI for Beginners: Where Should You Start?

A practical introduction to local AI, from choosing a runtime and model to testing quality on your own computer.

Use reference ↗

Category images: Unsplash

READING ORDER

Follow the setup
in sequence.

  1. 01

    Inventory the hardware

    Record memory, power, cooling, networking, and when each PC is available.

  2. 02

    Run the first model

    Establish a repeatable baseline with one runtime and one small model.

  3. 03

    Connect the devices

    Separate localhost and network problems, then restrict access to approved machines.

  4. 04

    Assign the work

    Split interactive and background roles, priorities, and ownership.

  5. 05

    Drill recovery

    Create reboots and disconnects deliberately, then verify requests and results survive.

“

We document what is reproducible, not merely fast.

Every guide includes its verification scope, update date, and primary sources.