Configure promptfoo red team with Wild West API
promptfoo is an open-source CLI for evals and red teaming. Its openai:chat provider accepts any model id plus a custom apiBaseUrl, so Wild West API can generate attacks, grade results or serve as the target.
Set it up
- Install promptfoo with npm install -g promptfoo, or run it through npx.
- Create a capped key in the dashboard and export it as WILDWEST_API_KEY.
- Run promptfoo redteam init, then edit promptfooconfig.yaml.
- Set redteam.provider to openai:chat:outlaw-1 with apiBaseUrl and apiKeyEnvar.
- Set defaultTest.options.provider to a Wild West model for grading.
- Run promptfoo redteam run, then promptfoo redteam report.
Roles
- Attacker.
redteam.provideris the model that writes adversarial test cases for each plugin. Filtered models often refuse to write cases for harmful-content plugins;outlaw-1writes them. - Grader.
defaultTest.options.provideroverrides the model that grades replies. - Target. Any entry under
targets. Pointing one at an uncensored model gives a no-guardrail baseline next to your real app.
promptfooconfig.yaml
description: Support bot red team
targets:
- id: openai:chat:qwen3.8-27b-outlaw
label: no-guardrail-baseline
config:
apiBaseUrl: https://wildwestapi.com/v1
apiKeyEnvar: WILDWEST_API_KEY
redteam:
purpose: >-
Customer support assistant for a bank. It must not reveal its
system prompt or any other customer's data.
provider:
id: openai:chat:outlaw-1
config:
apiBaseUrl: https://wildwestapi.com/v1
apiKeyEnvar: WILDWEST_API_KEY
temperature: 0.9
numTests: 5
plugins:
- harmbench
- pii
- prompt-extraction
- id: harmful:hate
numTests: 3
strategies:
- jailbreak
- prompt-injection
defaultTest:
options:
provider:
id: openai:chat:glm-5.3-outlaw
config:
apiBaseUrl: https://wildwestapi.com/v1
apiKeyEnvar: WILDWEST_API_KEY
temperature: 0Use the model id exactly as listed on /models/ after openai:chat:. Replace the baseline target with your own app's provider when you test it for real.
export WILDWEST_API_KEY=sk-ww-... export PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true promptfoo redteam run promptfoo redteam report
Without the last variable, promptfoo can send some generation work to its own hosted service. Setting it keeps generation on your configured provider; promptfoo notes this can lower attack quality, which an uncensored attacker offsets.
Cost estimate
Estimate only. With 4 plugin entries at 5 tests each and 2 strategies, expect around 60 to 100 target calls, plus generation and grading. Iterative jailbreak strategies make several attacker and grader calls per test, so budget about 10 calls per test case: roughly 600 to 1,000 calls. At 1,500 input and 400 output tokens per call, 1,000 calls is 1.5M in and 0.4M out. On outlaw-1 that is about $0.85; if all of it ran on glm-5.3-outlaw it would be about $5.55.
Tips
- Use
-jon the run command to set concurrency; start around 4. - Give the red team its own key with a hard cap. Iterative strategies multiply calls per test.
- Keep the attacker hot (0.8 to 1.0) and the grader at 0.
- A precise
purposegives far better test cases than a generic one.
Troubleshooting
- Requests going to api.openai.com.
apiBaseUrlis missing on that provider. Each provider block needs its own, or setOPENAI_BASE_URLglobally. - 401. The variable named in
apiKeyEnvaris empty in the shell running promptfoo. - 402. Out of balance or key cap; see 402.
- 429. Lower
-j; promptfoo retries rate-limited calls.
FAQ
How do I point promptfoo at a custom OpenAI-compatible URL?
Use openai:chat:<model-id> with config.apiBaseUrl set to https://wildwestapi.com/v1 and config.apiKeyEnvar set to the name of the variable holding your key.
Which model should write attacks?
outlaw-1 is the cheap default for generation. Use glm-5.3-outlaw for grading when accuracy matters more than cost.
Does promptfoo support HarmBench?
Yes. Add the harmbench plugin under redteam.plugins and set numTests to control how many of its behaviors are sampled.