VERIFICATIONChecked against current official documentation on 2026.08.04; hardware-specific performance is not generalized.

30-SECOND SUMMARY

What to take away

  • Measure load, first output, and sustained generation separately.
  • Repeat with fixed prompt, context, and output length.
  • Record worst cases, memory, failures, and quality.
BENCHMARK 01

Split perceived speed

Separate cold and warm runs.

  1. 01
    LOAD

    Model loading

  2. 02
    TTFT

    First output

  3. 03
    TOKENS

    Sustained generation

  4. 04
    QUALITY

    Task success

Use it this way Record median, worst case, memory, and failures.
SECTION 01

Separate three phases

Cold start includes model loading. Time to first token describes interactive waiting, while tokens per second describes continued generation.

The working rule for “Separate three phases” is: Measure load, first output, and sustained generation separately. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.

For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.

SECTION 02

Fix the conditions

Record model, quantization, runtime, context, prompt, output cap, power mode, and background workload. Separate cold and warm runs.

The working rule for “Fix the conditions” is: Repeat with fixed prompt, context, and output length. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.

Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.

SECTION 03

Read API metrics

Ollama generate and chat responses expose total, load, prompt-evaluation, and generation fields. Check units and the final streaming object.

The working rule for “Read API metrics” is: Record worst cases, memory, failures, and quality. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.

One successful run is not enough: repeat it after a restart and send one invalid input to confirm a controlled failure. Before connecting production data, test timeouts and cleanup so that an interrupted command does not leave stale processes, files, or application state.

curl http://localhost:11434/api/generate -d '{"model":"gemma3:4b","prompt":"Explain local AI","stream":false}'
SECTION 04

Pair speed with success

A fast wrong answer is not a better model. Compare performance only among candidates that pass the task-quality threshold.

The working rule for “Pair speed with success” is: Measure load, first output, and sustained generation separately. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.

Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.

PhaseMeaning
LoadMove model into memory
TTFTRequest to first output
GenerationOutput after first token
TotalEnd-to-end task time
FAQ

Frequently asked questions

Is tokens per second enough?

No. Interactive and batch work value different phases. For a practical check, follow the “Separate three phases” section, change one condition at a time, and record the result.

Can I measure once?

Repeat and record median plus worst case. Repeat with fixed prompt, context, and output length. For a practical check, follow the “Fix the conditions” section, change one condition at a time, and record the result.

Can models with different outputs be compared?

Use output limits and include quality scores. Record worst cases, memory, failures, and quality. For a practical check, follow the “Read API metrics” section, change one condition at a time, and record the result.

Primary sources

Check the original documentation for version-specific details.

Ollama Generate API Ollama Chat API llama.cpp repository

READ NEXT

How to Evaluate a Local LLM for Your Real WorkLocal AI for Beginners: Where Should You Start?