Wild West API

Run HarmBench against Wild West API

HarmBench is the Center for AI Safety's standardized benchmark for automated red teaming and refusal. It ships a behavior dataset, attack methods and a fine-tuned classifier, but its API model loader only recognizes a few vendor name patterns, so an OpenAI-compatible endpoint needs a small patch.

Set it up

  1. Clone centerforaisafety/HarmBench and install its requirements.
  2. Add a branch to api_models_map in api_models.py so Wild West model ids load the GPT class.
  3. Export OPENAI_BASE_URL=https://wildwestapi.com/v1 so the OpenAI client inside that class uses it.
  4. Add a closed_source entry for the model to configs/model_configs/models.yaml.
  5. Run scripts/run_pipeline.py with DirectRequest or HumanJailbreaks and --mode local.
  6. Run the evaluation step on a machine with a GPU for the HarmBench classifier.

Roles

Wild West API fits HarmBench mainly as the target. An uncensored model gives the high-compliance end of the scale, which helps you read attack success rates for your own model in context. HarmBench's own judge is the fine-tuned classifier cais/HarmBench-Llama-2-13b-cls, run locally, so the official scores do not use an API judge. Behavior categories include chemical and biological, cybercrime, harassment, misinformation and illegal activities; this page names them only.

The patch

api_models_map routes on substrings like gpt- and claude- and returns None for anything else. The GPT class builds OpenAI(api_key=..., timeout=...) with no base URL argument, but the OpenAI Python SDK falls back to OPENAI_BASE_URL.

# api_models.py, inside api_models_map, before `return None`
    elif model_name_or_path == 'outlaw-1' or model_name_or_path.endswith('-outlaw'):
        return GPT(model_name_or_path, token)
# configs/model_configs/models.yaml
outlaw_1:
  model:
    model_name_or_path: outlaw-1
    token: sk-ww-REPLACE_ME
  model_type: closed_source
export OPENAI_BASE_URL=https://wildwestapi.com/v1
python ./scripts/run_pipeline.py --methods DirectRequest,HumanJailbreaks \
  --models outlaw_1 --step all --mode local

Note that OPENAI_BASE_URL also redirects any real GPT entries in the same shell. The models.yaml token is a literal, so keep that file out of version control.

Cost estimate

Estimate only. The text behavior set is a few hundred behaviors. DirectRequest sends one completion each; HumanJailbreaks multiplies that by its template count. For 400 DirectRequest calls at about 100 tokens in and 512 out (the pipeline default --max_new_tokens), that is 40K in and 205K out: about $0.22 on outlaw-1. Template-based methods cost many times more because of long prompts. The classifier runs on your GPU, so it adds no API cost.

Tips

  • Start with --behavior_ids_subset on a handful of ids to check wiring.
  • Optimization methods such as GCG need model weights and do not apply to an API target. PAIR and TAP use their own attacker and judge configs, which need further code changes to use an API model.
  • If you only want the behaviors, the promptfoo harmbench plugin runs them with an API grader and no patching; see promptfoo.
  • Cap the key; long template prompts make HumanJailbreaks the expensive step.

Troubleshooting

  • Model returns None or a load error. The patch did not match the id; print model_name_or_path.
  • Every output is $ERROR$. The GPT class swallows API errors after 5 retries. Check the printed exception: 401 is a bad token, 402 is balance or cap.
  • Calls going to OpenAI. OPENAI_BASE_URL is not exported in that shell.

FAQ

Can HarmBench call an OpenAI-compatible endpoint out of the box?

Not for arbitrary model names. Its loader only matches certain vendor patterns, so you add one branch in api_models.py and set OPENAI_BASE_URL.

Can Wild West API replace the HarmBench classifier?

Not if you want comparable published numbers. The official metric uses the fine-tuned HarmBench classifier; an API judge gives a different, non-comparable score.

Is there an easier way to use HarmBench prompts?

Yes. promptfoo's harmbench plugin samples the behaviors and grades with a provider you configure.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.