Test LLM Guard with Wild West API behind it
LLM Guard is Protect AI's open-source library of input and output scanners for LLM apps. The scanners run locally; Wild West API supplies the model in between, and because it does not refuse, the output scanners are tested on their own merit.
Set it up
- Install with pip install llm-guard openai.
- Create a capped key and export it as WILDWEST_API_KEY.
- Point an OpenAI client at https://wildwestapi.com/v1.
- Choose input scanners such as PromptInjection and Toxicity, and output scanners such as NoRefusal and Sensitive.
- Run scan_prompt, call the model, then run scan_output on the reply.
- Log which scanner fired for each test prompt and compare it with your expected labels.
Why put an uncensored model behind a guardrail
With a safety-tuned model behind the guardrail, many harmful requests are refused by the model itself, so you never learn whether the output scanners would have caught the reply. With outlaw-1 behind it, the only thing between a test prompt and a harmful reply is LLM Guard, which shows its real catch rate. Wild West API can also act as the attacker, generating varied injection-style test prompts to measure the PromptInjection scanner's recall.
Code
import os
from openai import OpenAI
from llm_guard import scan_output, scan_prompt
from llm_guard.input_scanners import PromptInjection, TokenLimit, Toxicity
from llm_guard.output_scanners import NoRefusal, Sensitive
from llm_guard.output_scanners import Toxicity as OutputToxicity
client = OpenAI(api_key=os.environ["WILDWEST_API_KEY"],
base_url="https://wildwestapi.com/v1")
input_scanners = [PromptInjection(), Toxicity(), TokenLimit()]
output_scanners = [NoRefusal(), Sensitive(), OutputToxicity()]
def guarded(prompt: str) -> dict:
clean, ok, scores = scan_prompt(input_scanners, prompt)
if not all(ok.values()):
return {"blocked": "input", "scores": scores}
reply = client.chat.completions.create(
model="outlaw-1",
messages=[{"role": "user", "content": clean}],
temperature=0,
max_tokens=512,
).choices[0].message.content
out, ok, scores = scan_output(output_scanners, clean, reply)
if not all(ok.values()):
return {"blocked": "output", "scores": scores}
return {"blocked": None, "reply": out}Feed guarded() a labeled test set, for example prompts drawn from promptfoo or Spikee datasets, and count how often each layer blocks what it should.
Cost estimate
Estimate only. API cost covers only the prompts that pass the input scanners. For 1,000 test prompts with 70 percent passing, 700 calls at 200 tokens in and 400 out is 0.14M in and 0.28M out: about $0.32 on outlaw-1. The scanners run on your own CPU or GPU; the first run downloads their Hugging Face models.
Tips
- Use temperature 0 on the model so a scanner result is repeatable.
- The
NoRefusaloutput scanner flags refusals; with an uncensored model it should rarely fire, which is itself a check that the setup is right. - Scanners are the slow part. Batch test prompts and keep API concurrency modest.
- Use a capped key per test run.
Troubleshooting
- Name clash on Toxicity. Input and output scanners share names; import one under an alias as shown.
- 401. Key missing or not an
sk-ww-key. - 402. See 402 Payment Required.
- Slow first run. Scanner models are downloading; later runs use the cache.
FAQ
Does LLM Guard call an LLM itself?
No. Its scanners are local classifiers and rules. The LLM is whatever your app calls between scan_prompt and scan_output.
Why test with a model that does not refuse?
So the guardrail is measured alone. A refusing model hides output scanner misses.
Can I use the Anthropic endpoint instead?
Yes. Wild West API also serves /v1/messages, so an Anthropic SDK client works in the same spot.