Wild West API

Use Wild West API as Giskard's scan generator

Giskard is an open-source testing library for LLM apps and agents. Its v3 vulnerability_scan uses a generator model to write adversarial scenarios and judge the replies, and that model can be served by Wild West API.

Set it up

  1. Use Python 3.12 or newer, which Giskard v3 requires.
  2. Install with pip install "giskard[scan, litellm]".
  3. Create a capped key and export it as OPENAI_API_KEY in the scan's environment.
  4. Export OPENAI_BASE_URL and OPENAI_API_BASE as https://wildwestapi.com/v1.
  5. Call set_default_generator("openai/glm-5.3-flash-outlaw").
  6. Wrap your app as an async target function and run vulnerability_scan.

Which Giskard version

Giskard v3 is a rewrite split into giskard-checks and a scan package; the older v2 giskard.scan(model) API is still installable but no longer actively maintained. This page covers v3. The scan generates scenarios for a description of your app, sends them to your target, and judges the replies, in single-turn or multi-turn mode.

Roles

The default generator does both attack writing and judging. Filtered generators sometimes return weak or refused scenarios for security categories; an uncensored generator writes the full set and grades harmful replies without refusing. The target is your own async function, which can also call Wild West API for a no-guardrail baseline.

Code

The generator string is routed by provider prefix. Setting both OPENAI_BASE_URL (read by the OpenAI SDK) and OPENAI_API_BASE (read by LiteLLM) covers either backend.

export OPENAI_API_KEY="$WILDWEST_API_KEY"
export OPENAI_BASE_URL=https://wildwestapi.com/v1
export OPENAI_API_BASE=https://wildwestapi.com/v1
import asyncio, os
from openai import AsyncOpenAI
from pydantic import BaseModel
from giskard.checks import set_default_generator
from giskard.scan import vulnerability_scan

set_default_generator("openai/glm-5.3-flash-outlaw")

client = AsyncOpenAI()  # replace with your app's own client

class AgentInput(BaseModel):
    question: str

class AgentOutput(BaseModel):
    answer: str

async def support_bot(inputs: AgentInput) -> AgentOutput:
    r = await client.chat.completions.create(
        model="outlaw-1",
        messages=[
            {"role": "system", "content": "You are a bank support assistant."},
            {"role": "user", "content": inputs.question},
        ],
    )
    return AgentOutput(answer=r.choices[0].message.content)

result = asyncio.run(vulnerability_scan(
    target=support_bot,
    description="Bank support assistant. Must not reveal its system "
                "prompt or discuss other customers' accounts.",
    languages=["en"],
    max_scenarios=10,
))
print("passed:", result.passed_count, "failed:", result.failed_count)

Cost estimate

Estimate only. With max_scenarios=10 in single-turn mode, assume about 10 generation calls, 10 target calls and 10 judge calls, each around 1,500 tokens in and 500 out: 45K in and 15K out. On glm-5.3-flash-outlaw that is under $0.05. Multi-turn mode and larger scenario counts scale roughly linearly.

Tips

  • The description drives scenario generation, so state what the app must never do.
  • Start with max_scenarios=4 to check wiring, then raise it.
  • Switch the generator to glm-5.3-outlaw if verdicts look noisy.
  • Give the scan its own capped key; 402 means the cap or balance was hit.

Troubleshooting

  • Calls reach api.openai.com. The base URL variable was not set in the same process; export it before starting Python.
  • Model not found. Keep the openai/ prefix and use an exact id from /models/.
  • 401. OPENAI_API_KEY must hold the sk-ww- key for this process.

FAQ

Does Giskard work with OpenAI-compatible APIs?

Yes. Use the openai/ prefix on the model id and point the base URL at the compatible endpoint, here https://wildwestapi.com/v1.

Should I use Giskard v2 or v3?

v3 is the maintained line and the one this guide covers. v2 still installs but no longer gets active development.

Which Wild West model suits the generator?

glm-5.3-flash-outlaw is a good balance of cost and judging quality; glm-5.3-outlaw is the stronger option for final runs.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.