Wild West API

Vision models: sending images to an LLM

A vision model is a language model that can take images as input along with text and describe, reason about, or respond to what it sees.

What a vision model is

Vision-language models pair a language model with an image encoder. The encoder turns an image into a set of vectors that the language model reads as if they were extra tokens. The model then writes text as usual, so it can describe a picture, read text in a screenshot, or react in character to an image you send. It does not generate images; that is a separate kind of model.

How image input works in the API

In the OpenAI-style chat completions format, a message's content becomes an array of parts instead of a string:

{"model": "outlaw-1", "messages": [{
  "role": "user",
  "content": [
    {"type": "text", "text": "Describe this place in Mara's voice."},
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/..."}}
  ]
}]}

The image can be a base64 data URL or, where the server allows it, a public HTTPS URL. The Anthropic-style /v1/messages format uses an image content block with a base64 or URL source instead.

Cost and limits

Images are converted into tokens, and those tokens are billed as input. A single image can cost hundreds to a few thousand tokens depending on resolution and the model's encoder. In a chat, an image stays in the history and is resent with every later request until it falls out of the context, so it keeps costing. Many models downscale large images, so very high resolution rarely helps and only adds upload size.

Vision in roleplay

  • Sending a picture of a location, outfit or item for the character to react to.
  • Showing the model a character reference image.
  • SillyTavern's "Send inline images" option in Chat Completion lets you attach images to messages directly. The Image Captioning extension can instead turn an image into a text description, which works even with text-only models.

Common mistakes: sending images to a text-only model (the request fails or the image is dropped), leaving large images in a long chat, and expecting precise reading of tiny text.

Vision on Wild West API

outlaw-1, glm-5.3-flash-outlaw and mimo-v2.6-flash-outlaw accept image input. glm-5.3-outlaw is text only. Check the models page for the current list before building on a specific model. Use the image content format of whichever endpoint you call, and remember that browser-only apps cannot call the API directly because it does not answer CORS today.

FAQ

Can a vision model generate images?

No. Vision models read images and write text. Generating images needs a separate image model.

Do images cost tokens?

Yes. Each image is converted to input tokens, and it is billed again on every request while it stays in the chat history.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.