Run Inspect AI safety evals on Wild West API
Inspect is the UK AI Security Institute's open-source eval framework, and inspect_evals packages ready-made safety benchmarks for it. Its openai-api provider reads a base URL and key from environment variables named after any provider label you choose.
Set it up
- Install with pip install inspect-ai openai inspect-evals.
- Create a capped key in the dashboard.
- Export WILDWEST_API_KEY and WILDWEST_BASE_URL=https://wildwestapi.com/v1.
- Run inspect eval with --model openai-api/wildwest/<model-id>.
- Point the grader or judge at Wild West API with --model-role or a task -T parameter.
- Open results with inspect view.
How the provider naming works
Inspect accepts openai-api/<provider-name>/<model-name> and reads <PROVIDER_NAME>_API_KEY and <PROVIDER_NAME>_BASE_URL. With the label wildwest, those are exactly WILDWEST_API_KEY and WILDWEST_BASE_URL.
export WILDWEST_API_KEY=sk-ww-... export WILDWEST_BASE_URL=https://wildwestapi.com/v1 inspect eval inspect_evals/strong_reject \ --model openai-api/wildwest/qwen3.8-27b-outlaw \ -T judge_llm=openai-api/wildwest/glm-5.3-outlaw \ --max-connections 8 --limit 50
Roles
- Target. Running StrongREJECT or AgentHarm against an uncensored model gives the no-refusal end of the scale, useful when calibrating results for your own model.
- Judge. StrongREJECT's
judge_llmand AgentHarm'srefusal_judgeandsemantic_judgedefault to OpenAI models. Swapping inglm-5.3-outlawgives a judge that reads harmful replies without refusing. - Grader for your own tasks. Built-in scorers such as
model_graded_qa()use thegraderrole by default.
inspect eval my_task.py \ --model openai/gpt-4o-mini \ --model-role grader=openai-api/wildwest/glm-5.3-outlaw
AgentHarm runs the same way with -T refusal_judge=... and -T semantic_judge=.... All Wild West models support tool calling, which agent evals need.
Cost estimate
Estimate only. StrongREJECT's dataset is a few hundred forbidden prompts. With --limit 50: 50 target calls at about 150 in and 500 out, plus 50 judge calls at about 1,000 in and 200 out. Target on qwen3.8-27b-outlaw: about $0.06. Judge on glm-5.3-outlaw: about $0.13 + $0.05. Under $0.30 for the sample, and a few dollars for the full set. See pricing.
Tips
--max-connectionsdefaults to 32 per model. Start at 8 and raise it if no 429s appear.- Use
--limitfor a smoke test before a full run. - Keep the judge at temperature 0; target sampling can follow the eval's defaults or
--temperature. - One capped key per eval run makes cost per benchmark easy to read in the dashboard.
Troubleshooting
- Missing environment variable error. The label in the model string must match the variable prefix:
wildwestneedsWILDWEST_*. Hyphens in labels become underscores. - 401. Key unset in the shell running inspect.
- 402. See 402 Payment Required.
- Unknown model. The part after
wildwest/must match a listed id.
FAQ
What model string does Inspect use for Wild West API?
openai-api/wildwest/<model-id>, for example openai-api/wildwest/outlaw-1, with WILDWEST_API_KEY and WILDWEST_BASE_URL set.
How do I change the judge in an inspect_evals benchmark?
Pass the task's own parameter with -T, such as -T judge_llm=... for StrongREJECT, or use --model-role grader=... for scorers that read the grader role.
Is Inspect only for red teaming?
No. It is a general eval framework; the safety benchmarks in inspect_evals are one part of it.