30-SECOND SUMMARY
What to take away
- Measure load, first output, and sustained generation separately.
- Repeat with fixed prompt, context, and output length.
- Record worst cases, memory, failures, and quality.
Split perceived speed
Separate cold and warm runs.
- 01LOAD
Model loading
- 02TTFT
First output
- 03TOKENS
Sustained generation
- 04QUALITY
Task success
Separate three phases
Cold start includes model loading. Time to first token describes interactive waiting, while tokens per second describes continued generation.
The working rule for “Separate three phases” is: Measure load, first output, and sustained generation separately. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Fix the conditions
Record model, quantization, runtime, context, prompt, output cap, power mode, and background workload. Separate cold and warm runs.
The working rule for “Fix the conditions” is: Repeat with fixed prompt, context, and output length. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Read API metrics
Ollama generate and chat responses expose total, load, prompt-evaluation, and generation fields. Check units and the final streaming object.
The working rule for “Read API metrics” is: Record worst cases, memory, failures, and quality. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.
One successful run is not enough: repeat it after a restart and send one invalid input to confirm a controlled failure. Before connecting production data, test timeouts and cleanup so that an interrupted command does not leave stale processes, files, or application state.
curl http://localhost:11434/api/generate -d '{"model":"gemma3:4b","prompt":"Explain local AI","stream":false}'Pair speed with success
A fast wrong answer is not a better model. Compare performance only among candidates that pass the task-quality threshold.
The working rule for “Pair speed with success” is: Measure load, first output, and sustained generation separately. Use representative inputs with known answers and score accuracy, latency, memory, and consistency separately.
Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.
| Phase | Meaning |
|---|---|
| Load | Move model into memory |
| TTFT | Request to first output |
| Generation | Output after first token |
| Total | End-to-end task time |
Frequently asked questions
Is tokens per second enough?
No. Interactive and batch work value different phases. For a practical check, follow the “Separate three phases” section, change one condition at a time, and record the result.
Can I measure once?
Repeat and record median plus worst case. Repeat with fixed prompt, context, and output length. For a practical check, follow the “Fix the conditions” section, change one condition at a time, and record the result.
Can models with different outputs be compared?
Use output limits and include quality scores. Record worst cases, memory, failures, and quality. For a practical check, follow the “Read API metrics” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Generate API Ollama Chat API llama.cpp repository