Wild West API

Run Inspect AI safety evals on Wild West API

Inspect is the UK AI Security Institute's open-source eval framework, and inspect_evals packages ready-made safety benchmarks for it. Its openai-api provider reads a base URL and key from environment variables named after any provider label you choose.

Set it up

  1. Install with pip install inspect-ai openai inspect-evals.
  2. Create a capped key in the dashboard.
  3. Export WILDWEST_API_KEY and WILDWEST_BASE_URL=https://wildwestapi.com/v1.
  4. Run inspect eval with --model openai-api/wildwest/<model-id>.
  5. Point the grader or judge at Wild West API with --model-role or a task -T parameter.
  6. Open results with inspect view.

How the provider naming works

Inspect accepts openai-api/<provider-name>/<model-name> and reads <PROVIDER_NAME>_API_KEY and <PROVIDER_NAME>_BASE_URL. With the label wildwest, those are exactly WILDWEST_API_KEY and WILDWEST_BASE_URL.

export WILDWEST_API_KEY=sk-ww-...
export WILDWEST_BASE_URL=https://wildwestapi.com/v1

inspect eval inspect_evals/strong_reject \
  --model openai-api/wildwest/qwen3.8-27b-outlaw \
  -T judge_llm=openai-api/wildwest/glm-5.3-outlaw \
  --max-connections 8 --limit 50

Roles

  • Target. Running StrongREJECT or AgentHarm against an uncensored model gives the no-refusal end of the scale, useful when calibrating results for your own model.
  • Judge. StrongREJECT's judge_llm and AgentHarm's refusal_judge and semantic_judge default to OpenAI models. Swapping in glm-5.3-outlaw gives a judge that reads harmful replies without refusing.
  • Grader for your own tasks. Built-in scorers such as model_graded_qa() use the grader role by default.
inspect eval my_task.py \
  --model openai/gpt-4o-mini \
  --model-role grader=openai-api/wildwest/glm-5.3-outlaw

AgentHarm runs the same way with -T refusal_judge=... and -T semantic_judge=.... All Wild West models support tool calling, which agent evals need.

Cost estimate

Estimate only. StrongREJECT's dataset is a few hundred forbidden prompts. With --limit 50: 50 target calls at about 150 in and 500 out, plus 50 judge calls at about 1,000 in and 200 out. Target on qwen3.8-27b-outlaw: about $0.06. Judge on glm-5.3-outlaw: about $0.13 + $0.05. Under $0.30 for the sample, and a few dollars for the full set. See pricing.

Tips

  • --max-connections defaults to 32 per model. Start at 8 and raise it if no 429s appear.
  • Use --limit for a smoke test before a full run.
  • Keep the judge at temperature 0; target sampling can follow the eval's defaults or --temperature.
  • One capped key per eval run makes cost per benchmark easy to read in the dashboard.

Troubleshooting

  • Missing environment variable error. The label in the model string must match the variable prefix: wildwest needs WILDWEST_*. Hyphens in labels become underscores.
  • 401. Key unset in the shell running inspect.
  • 402. See 402 Payment Required.
  • Unknown model. The part after wildwest/ must match a listed id.

FAQ

What model string does Inspect use for Wild West API?

openai-api/wildwest/<model-id>, for example openai-api/wildwest/outlaw-1, with WILDWEST_API_KEY and WILDWEST_BASE_URL set.

How do I change the judge in an inspect_evals benchmark?

Pass the task's own parameter with -T, such as -T judge_llm=... for StrongREJECT, or use --model-role grader=... for scorers that read the grader role.

Is Inspect only for red teaming?

No. It is a general eval framework; the safety benchmarks in inspect_evals are one part of it.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.