VERIFICATIONFocuses on a reproducible evaluation process and does not present invented benchmark results.

30-SECOND SUMMARY

What to take away

  • Create 10–20 representative prompts with known evidence.
  • Change one variable at a time and repeat each prompt.
  • Track critical failures and operational limits, not only average quality.
EVALUATION 01

A repeatable model-selection loop

Change one variable per comparison.

  1. 01
    PROMPTS

    Fixed cases

  2. 02
    REPEAT

    Run multiple times

  3. 03
    SCORE

    Quality and operations

  4. 04
    DECIDE

    Apply thresholds

Use it this way Record worst cases and critical failures, not only averages.
SECTION 01

Define the job

Replace ‘find the smartest model’ with a concrete objective such as ‘summarize English support tickets without inventing refund promises.’

Specify the device, acceptable latency, memory ceiling, language, and disqualifying failures.

For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.

SECTION 02

Build a small representative set

Include common requests, rare exceptions, formatting, answer-not-found, conflicting instructions, and safety cases. Attach expected answer points and source evidence.

Use public or anonymized data so the evaluation set does not become a new sensitive repository.

Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.

SECTION 03

Fix the conditions

Record prompt, temperature, context, runtime version, quantization, and hardware state. Change only one variable per comparison.

Separate first-run loading time from repeated generation and run each prompt at least three times to observe variance.

Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.

SECTION 04

Score quality and operations separately

Quality can include factual accuracy, instruction following, evidence alignment, format validity, and language quality. Operations can include time to first output, total generation time, peak memory, and failure rate.

A critical error may disqualify a model even if its average score is high.

Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.

TestWhat to recordCritical failure example
Fact extractionCorrect names and numbersInvented value
SummaryCoverage and additionsUnsupported claim
Structured outputSchema validityParsing failure
No answerAppropriate uncertaintyConfident fabrication
OperationsLatency, memory, errorsCrash or timeout
SECTION 05

Switch models only when the gain matters

Define a threshold that a new model must exceed before replacing the current one. This prevents constant churn whenever a new release appears.

Rerun the same set after runtime or model updates and keep a rollback path for regressions.

Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.

FAQ

Frequently asked questions

Are public benchmarks useless?

No. They help shortlist candidates, but they do not replace tests on your language, documents, hardware, and failure conditions. For a practical check, follow the “Define the job” section, change one condition at a time, and record the result.

How many prompts should I start with?

A carefully reviewed set of 10–20 is practical. Add real failures as permanent regression cases. For a practical check, follow the “Build a small representative set” section, change one condition at a time, and record the result.

What speed metric matters most?

Measure time to first output and sustained generation separately, and distinguish first from repeated runs. Track critical failures and operational limits, not only average quality. For a practical check, follow the “Fix the conditions” section, change one condition at a time, and record the result.

Primary sources

Check the original documentation for version-specific details.

Ollama Chat API Metrics LM Studio Model Download Guide

READ NEXT

Benchmark Local LLM Speed CorrectlyLocal AI for Beginners: Where Should You Start?