Run garak against Wild West API as target or judge
garak is NVIDIA's open-source LLM vulnerability scanner. It fires probe prompts at a model and runs detectors over the replies, and it talks to any OpenAI-compatible endpoint through its openai.OpenAICompatible generator.
Set it up
- Install garak with pip install -U garak.
- Create a key in the dashboard and give it a hard spend cap sized for one scan.
- Export the key as OPENAICOMPATIBLE_API_KEY, which is the variable the OpenAICompatible generator reads.
- Save a YAML config that sets the generator uri to https://wildwestapi.com/v1/ and suppresses the default stop sequences.
- Run garak with --config, --target_type openai.OpenAICompatible, --target_name set to a model id and a --spec probe selection.
- Read the JSONL and HTML report paths garak prints when the run finishes.
What garak does and where Wild West API fits
garak ships probe modules such as dan, encoding, promptinject, latentinjection, leakreplay, malwaregen and packagehallucination. Each probe sends a batch of prompts, and detectors score the replies. Wild West API can sit in two places:
- Target. Scan
outlaw-1orqwen3.8-27b-outlawto get a no-guardrail baseline. Comparing that hit rate with your production model shows how much your filtering is actually doing. - Judge. The
judge.ModelAsJudge,judge.Refusalandjudge.Jailbreakdetectors call an OpenAI-format model to grade replies. A filtered judge sometimes refuses to read the harmful output it is meant to grade; an uncensored one does not.
Config file
Pass the endpoint in a YAML file. A GitHub issue reports that --generator_options with a uri was ignored for this generator, so the config file route is the reliable one. The generator defaults to stop: ["#", ";"], which would cut replies at the first semicolon, so suppress it.
# ww-garak.yaml
plugins:
generators:
openai:
OpenAICompatible:
uri: https://wildwestapi.com/v1/
suppressed_params:
- stop
- seed
- frequency_penalty
- presence_penalty
detectors:
judge:
detector_model_type: openai.OpenAICompatible
detector_model_name: glm-5.3-flash-outlaw
detector_model_config:
uri: https://wildwestapi.com/v1/Then run a scan:
export OPENAICOMPATIBLE_API_KEY="$WILDWEST_API_KEY" python -m garak --config ww-garak.yaml \ --target_type openai.OpenAICompatible \ --target_name outlaw-1 \ --spec probes.promptinject,probes.encoding \ --generations 2 --parallel_attempts 8
To grade with the judge, add --detectors judge.Jailbreak or judge.Refusal. --probes still works but the CLI marks it deprecated in favor of --spec.
Cost estimate
Rough estimate, not a quote. A focused scan of a few probe modules is often 2,000 to 5,000 prompts. With --generations 2, 5,000 prompts is 10,000 calls. At about 250 input and 300 output tokens each, that is 2.5M input and 3M output tokens. On outlaw-1 ($0.30 in, $1.00 out per million) that comes to roughly $0.75 + $3.00 = $3.75. Adding a glm-5.3-flash-outlaw judge pass over the same replies adds a similar amount. A full --spec all run is far larger, so cap the key first. See pricing.
Tips
- Raise
--parallel_attemptsfor API models; start at 8 and go up if you see no 429s. - Make one key per scan with its own spend cap, so a runaway
--spec allstops at a known number. - Keep
--generationslow (1 or 2) for a first pass; the default is 5. - Use tags to scope work, for example
--spec tag:owasp:llm01for prompt injection. Background: LLM01 prompt injection.
Troubleshooting
- Connection to localhost:8000. The uri was not picked up; that is the generator default. Check the YAML nesting under
plugins.generators.openai.OpenAICompatible. - 401.
OPENAICOMPATIBLE_API_KEYis unset or not ansk-ww-key. - 402. Balance or key cap reached; see 402 Payment Required.
- Unknown model. Use an id from /models/ exactly.
- 429 or 5xx. garak retries 408, 429, 502, 503 and 504 by default; lower parallelism if they persist.
FAQ
Which garak generator should I use for Wild West API?
Use openai.OpenAICompatible with uri set to https://wildwestapi.com/v1/ in a YAML config. The plain openai generator is meant for OpenAI's own service.
Can the judge detectors use an uncensored model?
Yes. The judge detectors accept any generator that is OpenAI-compatible, configured through detector_model_type, detector_model_name and detector_model_config.
Why scan an uncensored model at all?
It gives a baseline with no safety tuning. The gap between that hit rate and your production model's hit rate is a direct measure of what your guardrails add.