Wild West API

Configure promptfoo red team with Wild West API

promptfoo is an open-source CLI for evals and red teaming. Its openai:chat provider accepts any model id plus a custom apiBaseUrl, so Wild West API can generate attacks, grade results or serve as the target.

Set it up

  1. Install promptfoo with npm install -g promptfoo, or run it through npx.
  2. Create a capped key in the dashboard and export it as WILDWEST_API_KEY.
  3. Run promptfoo redteam init, then edit promptfooconfig.yaml.
  4. Set redteam.provider to openai:chat:outlaw-1 with apiBaseUrl and apiKeyEnvar.
  5. Set defaultTest.options.provider to a Wild West model for grading.
  6. Run promptfoo redteam run, then promptfoo redteam report.

Roles

  • Attacker. redteam.provider is the model that writes adversarial test cases for each plugin. Filtered models often refuse to write cases for harmful-content plugins; outlaw-1 writes them.
  • Grader. defaultTest.options.provider overrides the model that grades replies.
  • Target. Any entry under targets. Pointing one at an uncensored model gives a no-guardrail baseline next to your real app.

promptfooconfig.yaml

description: Support bot red team
targets:
  - id: openai:chat:qwen3.8-27b-outlaw
    label: no-guardrail-baseline
    config:
      apiBaseUrl: https://wildwestapi.com/v1
      apiKeyEnvar: WILDWEST_API_KEY

redteam:
  purpose: >-
    Customer support assistant for a bank. It must not reveal its
    system prompt or any other customer's data.
  provider:
    id: openai:chat:outlaw-1
    config:
      apiBaseUrl: https://wildwestapi.com/v1
      apiKeyEnvar: WILDWEST_API_KEY
      temperature: 0.9
  numTests: 5
  plugins:
    - harmbench
    - pii
    - prompt-extraction
    - id: harmful:hate
      numTests: 3
  strategies:
    - jailbreak
    - prompt-injection

defaultTest:
  options:
    provider:
      id: openai:chat:glm-5.3-outlaw
      config:
        apiBaseUrl: https://wildwestapi.com/v1
        apiKeyEnvar: WILDWEST_API_KEY
        temperature: 0

Use the model id exactly as listed on /models/ after openai:chat:. Replace the baseline target with your own app's provider when you test it for real.

export WILDWEST_API_KEY=sk-ww-...
export PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true
promptfoo redteam run
promptfoo redteam report

Without the last variable, promptfoo can send some generation work to its own hosted service. Setting it keeps generation on your configured provider; promptfoo notes this can lower attack quality, which an uncensored attacker offsets.

Cost estimate

Estimate only. With 4 plugin entries at 5 tests each and 2 strategies, expect around 60 to 100 target calls, plus generation and grading. Iterative jailbreak strategies make several attacker and grader calls per test, so budget about 10 calls per test case: roughly 600 to 1,000 calls. At 1,500 input and 400 output tokens per call, 1,000 calls is 1.5M in and 0.4M out. On outlaw-1 that is about $0.85; if all of it ran on glm-5.3-outlaw it would be about $5.55.

Tips

  • Use -j on the run command to set concurrency; start around 4.
  • Give the red team its own key with a hard cap. Iterative strategies multiply calls per test.
  • Keep the attacker hot (0.8 to 1.0) and the grader at 0.
  • A precise purpose gives far better test cases than a generic one.

Troubleshooting

  • Requests going to api.openai.com. apiBaseUrl is missing on that provider. Each provider block needs its own, or set OPENAI_BASE_URL globally.
  • 401. The variable named in apiKeyEnvar is empty in the shell running promptfoo.
  • 402. Out of balance or key cap; see 402.
  • 429. Lower -j; promptfoo retries rate-limited calls.

FAQ

How do I point promptfoo at a custom OpenAI-compatible URL?

Use openai:chat:<model-id> with config.apiBaseUrl set to https://wildwestapi.com/v1 and config.apiKeyEnvar set to the name of the variable holding your key.

Which model should write attacks?

outlaw-1 is the cheap default for generation. Use glm-5.3-outlaw for grading when accuracy matters more than cost.

Does promptfoo support HarmBench?

Yes. Add the harmbench plugin under redteam.plugins and set numTests to control how many of its behaviors are sampled.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.