30-SECOND SUMMARY
What to take away
- Create 10–20 representative prompts with known evidence.
- Change one variable at a time and repeat each prompt.
- Track critical failures and operational limits, not only average quality.
A repeatable model-selection loop
Change one variable per comparison.
- 01PROMPTS
Fixed cases
- 02REPEAT
Run multiple times
- 03SCORE
Quality and operations
- 04DECIDE
Apply thresholds
Define the job
Replace ‘find the smartest model’ with a concrete objective such as ‘summarize English support tickets without inventing refund promises.’
Specify the device, acceptable latency, memory ceiling, language, and disqualifying failures.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Build a small representative set
Include common requests, rare exceptions, formatting, answer-not-found, conflicting instructions, and safety cases. Attach expected answer points and source evidence.
Use public or anonymized data so the evaluation set does not become a new sensitive repository.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Fix the conditions
Record prompt, temperature, context, runtime version, quantization, and hardware state. Change only one variable per comparison.
Separate first-run loading time from repeated generation and run each prompt at least three times to observe variance.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
Score quality and operations separately
Quality can include factual accuracy, instruction following, evidence alignment, format validity, and language quality. Operations can include time to first output, total generation time, peak memory, and failure rate.
A critical error may disqualify a model even if its average score is high.
Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.
| Test | What to record | Critical failure example |
|---|---|---|
| Fact extraction | Correct names and numbers | Invented value |
| Summary | Coverage and additions | Unsupported claim |
| Structured output | Schema validity | Parsing failure |
| No answer | Appropriate uncertainty | Confident fabrication |
| Operations | Latency, memory, errors | Crash or timeout |
Switch models only when the gain matters
Define a threshold that a new model must exceed before replacing the current one. This prevents constant churn whenever a new release appears.
Rerun the same set after runtime or model updates and keep a rollback path for regressions.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Frequently asked questions
Are public benchmarks useless?
No. They help shortlist candidates, but they do not replace tests on your language, documents, hardware, and failure conditions. For a practical check, follow the “Define the job” section, change one condition at a time, and record the result.
How many prompts should I start with?
A carefully reviewed set of 10–20 is practical. Add real failures as permanent regression cases. For a practical check, follow the “Build a small representative set” section, change one condition at a time, and record the result.
What speed metric matters most?
Measure time to first output and sustained generation separately, and distinguish first from repeated runs. Track critical failures and operational limits, not only average quality. For a practical check, follow the “Fix the conditions” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Chat API Metrics LM Studio Model Download Guide