Stop Choosing Models by Feel: Turn LLM Evaluation, Red Teaming, and CI into a Repeatable Workflow with Promptfoo
Stop Choosing Models by Feel: Turn LLM Evaluation, Red Teaming, and CI into a Repeatable Workflow with Promptfoo
I think the first problem exposed when an LLM application moves from prototype to production is usually not that the model is insufficiently intelligent. It is that the team lacks a reliable way to answer three questions: did this change improve the answer, did it improve only a small set of examples, and did quality or security get worse as a result? If every change is tested by typing a few prompts, reading a few outputs, and deciding from memory whether the result feels better, the process quickly loses reproducibility.
Promptfoo is an open-source CLI and library for evaluating and red-teaming LLM applications. It brings prompts, providers, test data, assertions, and result inspection into one workflow. Its official documentation also covers CI/CD, code scanning, and vulnerability scanning. What makes it interesting to me is not that it is another prompt utility. It turns decisions scattered across chat windows, spreadsheets, and manual reviews into engineering assets that can be versioned, rerun, and checked in a Pull Request.
The short version: Promptfoo answers how to trust a change
If your team only compares two models occasionally, calling an API and reading the outputs manually may be enough. Once an application has multiple prompts, providers, fixed business cases, or frequent deployment changes, evaluation needs to become a regression test rather than a one-off experiment.
The basic unit of Promptfoo can be understood as four layers:
1. Prompts: the instructions being tested, stored inline or in files.
2. Providers: the models or LLM APIs being compared. The official examples include OpenAI providers, and the project documents integrations with more model services and local models.
3. Tests: input variables and expected conditions, so the same cases can run repeatedly.
4. Assertions: checks on the output, such as required text, structure, model-based grading, or security conditions.
The value of these layers is that the model is not the only thing being evaluated. Hold the test cases constant while comparing prompts, hold the prompt constant while comparing providers, and put quality checks and security checks into the same change process. When results get worse, there is at least a traceable difference instead of only the statement that the new version feels strange.
Why manual prompt testing breaks down
Manual testing is useful, but it is a poor sole quality gate. The first problem is case drift: one day you test three easy questions, then change the data source or language and continue relying on the old impression. The second is unfair comparison. If different models do not receive exactly the same prompt, parameters, and inputs, the apparent winner may simply reflect the test design.
The third problem is that negative cases are often ignored. When a user asks for a normal article summary, most models look good. Differences appear when an input contains an escalation attempt, sensitive data, prompt injection, or incomplete constraints. The fourth problem is release cadence. A manual check may discover a problem once, but it will not automatically repeat the same check on the next commit.
I would place Promptfoo between the model call and the product test suite. It does not replace integration tests, and it does not guarantee that a model will always be correct. It supplies an evaluation layer dedicated to the variability of LLM output. Its job is to let a team continuously observe quality, differences, and risks using one maintained set of cases.
Start with a minimal, verifiable eval
The official Quick Start requires Node.js >=22.22.0 and recommends Node.js 24 LTS. You can start with npx without installing the CLI globally; npm and Homebrew installation are also documented. Never put a real credential in a configuration file. Provider keys should come from environment variables or CI secrets.
npx promptfoo@latest init --example getting-started
cd getting-started
export OPENAI_API_KEY="$YOUR_PROVIDER_API_KEY"
npx promptfoo@latest eval
npx promptfoo@latest view
Each command has a clear responsibility. init creates an editable example directory, the environment variable enables the provider call, eval runs the evaluation, and view opens the result viewer. The first success criterion is not that the answer sounds human. It is that another clean environment can rerun the same setup and produce a result that can be compared.
You can then reduce the setup to a small promptfooconfig.yaml. This example follows the configuration concepts shown in the official documentation and compares two prompts against the same input:
description: customer support answer evaluation
prompts:
- file://prompt1.txt
- file://prompt2.txt
providers:
- openai:gpt-5-mini
tests:
- vars:
question: "What information is required for a refund?"
assert:
- type: contains
value: "refund"
Adjust the provider name according to the current provider documentation and the model you can actually use. Do not treat an example model name as a permanent product promise. Do not begin with one test case. I suggest three groups: common normal inputs, boundary inputs that commonly fail, and requests that the system must never execute or disclose. Even a few dozen cases are a better maintenance baseline than a single manual demonstration.
Assertions are more than keyword checks
The simplest assertion checks whether output contains or does not contain a string. That is useful for explicit product rules, such as mentioning a refund window or excluding an internal field name. But a keyword does not prove that the answer is correct. An output can contain three numbered items while still having the wrong order, incomplete information, or false claims.
I divide assertions into three levels. The first is deterministic checks: strings, JSON structure, regular expressions, length, and forbidden content. They are fast, inexpensive, and stable in CI. The second is semantic grading: a grader evaluates relevance, completeness, tone, or similarity to a reference answer. This handles open-ended answers better, but the grader has its own bias and variability. The third is domain review: people who understand the business periodically inspect cases and failures to confirm that the metrics still represent product quality.
These levels cannot replace one another. Keyword-only checks create false confidence, while judge-only evaluation delegates an unstable decision to another model. A practical approach is to run cheap, deterministic rules first, reserve expensive semantic checks for cases that need them, and use human sampling to calibrate the results.
Compare models by quality, cost, and failure mode
Promptfoo is well suited to side-by-side model comparison, but the highest score should not automatically become the production choice. At least four dimensions matter: quality, latency, cost, and failure mode.
Quality can combine pass rate, grader scores, and human samples. Latency needs tail behavior, not only the average. Cost includes prompt length, retry count, and output length rather than only the advertised unit price. Failure mode asks whether a model honestly says it does not know or fills the gap with fluent but incorrect text.
I would store these evaluation reports as versioned artifacts instead of relying on a dashboard screenshot. When a prompt or provider changes, record the test-data version, configuration, run date, and passing or failing cases. If a model scores slightly higher on average but has a worse error rate on sensitive cases, the difference should not be hidden by a flattering aggregate score.
Caching and nondeterminism also matter. The README highlights caching and live reload as developer-focused features, but caching changes which calls you observe, and stochastic output means repeated runs may differ. Your evaluation specification should say what can be reused, what must be rerun, and how borderline scores are handled. Repeatable does not mean every output is byte-for-byte identical; it means that the conditions, cases, and decision method are consistent.
Red teaming: put security checks in the same pipeline
Quality evaluation asks whether the system works under normal conditions. Red teaming asks whether it fails when deliberately challenged. Promptfoo provides official red-team and vulnerability-scanning documentation, and its CLI can run redteam run. Security testing can therefore become a scheduled check instead of a one-time attack demonstration.
npx promptfoo@latest redteam run --config promptfooconfig.yaml
Define the scope before scanning. Which data may never be disclosed? Which tool calls must be refused? Which permissions must not be rewritten by input text? Which output-format failure could cause a downstream system to act incorrectly? Without these policies, a scan easily becomes a long list of alarming findings that nobody can prioritize.
I would classify results into acceptable refusals, ambiguous cases requiring review, and high-risk cases that block release. For every high-risk case, keep a minimal reproduction input and a regression test after the fix. Use de-identified or synthetic data for sensitive tests, and make sure reports and CI logs do not print API keys, personal data, or complete confidential system prompts.
A red-team tool is not a security certificate. It expands attack coverage, but it cannot automatically understand your authorization model, retention policy, business impact, or the behavior of the database and tool-execution layers around the model. Security teams still need to review policies, severity, and remediation.
Connect it to CI/CD
Once the evaluation configuration and test data live in the repository, the next step is CI. The official CI/CD documentation shows npx promptfoo@latest eval with an output file and also documents redteam run. A minimal pipeline can look like this:
npx promptfoo@latest eval -c promptfooconfig.yaml -o results.json
npx promptfoo@latest redteam run promptfooconfig.yaml
Do not put every model, every case, and every red-team scenario into every Pull Request on day one. A more reliable split is a fast smoke eval on every PR, a full regression suite after merging or on a schedule, and expensive model comparisons or deep red-team scans at night or before release. This keeps feedback fast enough that the team will not bypass the checks.
Failure thresholds should match the test type. A deterministic assertion can block immediately. A semantic score near the threshold can request human review. A temporary provider outage should be reported separately from a real product regression. Otherwise a network error can be mistaken for a prompt-quality problem.
CI secrets belong only in the execution environment. They should never be committed to promptfooconfig.yaml, evaluation output, or issue comments. Before sharing results, check whether they contain user input, model output, or internal system prompts. The goal of automation is not to publish every result; it is to give the right people enough evidence with the right permissions.
Where not to overpromise
First, Promptfoo cannot define what a good answer means for your product. It provides execution and comparison, but test data, assertions, thresholds, and risk categories remain the team's responsibility. Poor test data only produces automated evidence for the wrong conclusion.
Second, an LLM judge is not perfectly objective. Its grading prompt, reference answer, model version, and context can all change the result. Keep deterministic checks and human samples for important features.
Third, providers and runtime conditions affect reproducibility. Model updates, service limits, network errors, rate limits, and price changes can make a run fail. Record the provider, model identifier, configuration, and runtime; use fixed versions or controlled retries when appropriate.
Fourth, a quick CLI setup is not complete governance. Production adoption also needs test-data versioning, output retention, access control, PII masking, secret management, failure triage, and change approval. These belong to the product lifecycle and do not appear automatically after installation.
How I would adopt it in the first week
On day one, choose one high-value flow such as customer-support summarization or knowledge-base Q&A. Create ten to twenty normal cases and five boundary cases. On day two, put prompts, providers, and test data in configuration and start with deterministic assertions. On day three, add one semantic quality check and compare its scores with expert judgment through manual sampling.
On day four, write three to five security policies and add a small red-team suite without trying to cover every attack. On day five, connect a smoke eval to Pull Requests and experience both a pass and a failure. After a failure, the team should be able to locate the case, understand the cause, change the configuration, and retain the regression case. The first-week success metric is not the number of tests. It is whether everyone begins discussing changes with the same evidence.
Add model comparisons, full CI, scheduled scans, and team reports gradually. As the suite grows, remove cases that no longer discriminate, merge duplicates, and review thresholds. Evaluation suites decay just like product code. Adding tests without maintaining them eventually creates another dashboard nobody trusts.
Conclusion: replace “feels better” with “enough evidence”
Promptfoo does not choose one model that is always best. Its core value is putting model and prompt changes into an inspectable engineering workflow. Start with npx promptfoo@latest init --example getting-started, establish a baseline with a small case set, and add assertions, model comparison, red teaming, and CI/CD one layer at a time.
I would recommend it to teams that already have an LLM application, frequently change prompts, or need to choose among providers. If you do not yet have stable test data or a definition of product quality, do not rush to install it. Write the success conditions and security policies first so the tool has something meaningful to execute. Mature practice is not chasing a beautiful score. It is being able to explain after every change what was tested, what improved, what regressed, and why the release is still safe enough.