Why one impressive answer is not enough

When choosing a local model for a knowledge base, customer support, or document processing, it is easy to run five attractive demos and mistake them for validation. Real work brings outdated instructions, ambiguous product codes, irrelevant questions, and requests for facts absent from the source material. Compare systems on the same pre-labeled cases, not on the impression left by a conversation.

Promptfoo is an open-source tool for this purpose. Its documentation covers test sets with variables and assertions, side-by-side provider comparisons, result review, and connections to local models through Ollama or vLLM's OpenAI-compatible interface. The official product screenshot shows a matrix: queries in rows, model or prompt variants in columns, and responses with test results in the cells. It illustrates the interface, not a test we ran or proof of any particular model's quality.

This is useful to small and medium-sized businesses once they have chosen a single process and need to decide whether a local model is ready for it. It is not a new model release. It is a practical way to avoid buying a server or granting an agent access based on three successful demonstration questions.

Define the decision before the metric

Start with the decision an employee needs to make, not with abstract intelligence. For an internal knowledge base, that may mean an answer citing the current document and a refusal when evidence is missing. For ticket routing, it may mean the right queue plus a flag for uncertainty. For extracting document fields, it means correct values and units, with an explicit indication when the source is insufficient.

A first pilot with 40–80 carefully selected cases can be more useful than a thousand unlabeled synthetic questions. Build it from anonymized real requests and previously observed failures. Cover four groups:

  • routine cases where the system should save work;
  • boundary cases: incomplete requests, mixed languages, or similar product names;
  • refusal cases: stale documents, missing facts, and requests beyond the user's access rights;
  • critical failures: invented prices, false promises to customers, or action without approval.

Each row needs more than an expected sentence. Record the source document and version, acceptable answer variants, prohibited claims, and the cost of a mistake. Sometimes “hand this to a person” is the only correct answer. If two employees cannot agree on the label, you cannot honestly score the model on that case: clarify the business rule first.

The NIST AI Risk Management Framework recommends measuring performance in the intended context of use and documenting differences between tests and the operating environment. A general public benchmark does not know your catalog, document-access rules, or the cost of a wrong decision.

Build a comparison matrix

Run two or three candidates against the same set: perhaps the current model and a local challenger, or one model with two prompt versions. Pin the model weights, chat template, system instructions, document set, generation settings, and retrieval configuration. Otherwise, an apparent model improvement could be caused by a refreshed index.

Separate checks by reliability. Deterministic rules work for a JSON schema, mandatory field, source citation, forbidden customer identifier, or exact arithmetic. Semantic grading is useful for meaning and completeness, but an automated model judge can be wrong too. A subject-matter expert should review a sample of its verdicts, especially for high-impact cases. One overall pass percentage is not a substitute for expert review.

Promptfoo's documentation supports checks such as exact match, regular expressions, structural formats, and model-based grading. Test data can live in local CSV, JSON, JSONL, or YAML files. Results can be exported to JSON and HTML for failure analysis. A small pilot needs only a versioned case file and a report; there is no need to buy an elaborate experiment-management platform.

Every pass rate needs a denominator. If there are ten critical cases, one wrong answer is already a 10% failure rate for that group, even if the overall score on easy questions looks attractive. Track correct answers, safe refusals, confident wrong answers, response time, and cost per accepted result separately. Include the time people spend correcting each candidate's mistakes.

Where data goes—and why that needs a separate check

Running the utility locally does not mean the entire evaluation stays inside the company. Requests go to the model provider configured for the test. A model judge has its own provider; Promptfoo documents external defaults and how to override them explicitly. Configure the intended local endpoint for every path—generation, grading, and embedding-based checks—or start with deterministic assertions and human review.

Promptfoo also has usage telemetry that can be disabled through a documented setting. The project says the telemetry does not contain prompt or response text, but the decision to send metadata still belongs to the company's policy. Do not share or publish reports containing operational data. Reports and exports can hold queries, variables, model responses, and configuration; treat them as internal documents, set retention periods, and remove personal data before testing.

A local model needs more than weights and a GPU. It needs a stable endpoint, a pinned model version, concurrency limits, and latency measurements under realistic load. A single request on a laptop says little about a queue of live business queries. If some checks need a network service or a model judge, include their compute and data paths in the architecture and budget.

The cost of testing

Pilot cost includes labeling cases, engineering the integration, running the candidates, subject-matter review of disputed responses, and updating the set as new failures appear. Local execution adds hardware and operating time; an API adds request charges. The important denominator is not generated tokens but answers accepted without dangerous correction.

An illustrative planning calculation: 60 test cases, two model options, and two prompt variants make 240 generations in a full run. If an expert spends two minutes manually reviewing each unique response, that is up to eight hours before preparation and dispute resolution. Automated assertions reduce review of obvious failures but do not eliminate scrutiny of critical answers. These are modeling assumptions, not Promptfoo measurements or universal labor benchmarks.

First establish whether the new configuration creates value in one process. If case volume is low and mistakes are expensive, preserving human decisions may be cheaper and safer than automation. If a local candidate wins on quality but needs an expensive always-on GPU for a few queries a day, the economic answer can still be no.

A practical next step

Assign a process owner and collect 40–80 anonymized cases from recent work. Label the correct result and critical prohibitions; pin the source-document versions. Compare two candidates on the same set, then inspect every failure and a random sample of apparent successes. After changes, repeat the evaluation on a held-out portion of cases that was not used to tune the prompt.

Define the operational gate before running the test: zero critical failures on the trial set, an acceptable safe-refusal rate, latency during busy hours, and cost per accepted answer. Zero failures in a small sample does not prove there will be none in production. Start by assisting a human and keep collecting real failures. A comparison matrix is not for a polished report; it makes every model, prompt, or retrieval change comparable with the previous version.

Image: official Promptfoo product screenshot, © Promptfoo 2025. The original file and MIT License terms are listed in the sources. The square cover is a crop with no generated extension.