Write YARA rules from your malware analysis notes
After you analyse a sample in your lab you have notes: a mutex name, a PDB path, a User-Agent string. Writing those into a YARA rule is routine work. Apex Flash can draft the rule from the notes; you then test it for misses and false positives.
Set it up
- Write down the indicators from your analysis: stable strings, byte patterns, file size range and format.
- Remove anything unique to a single build, such as a timestamp.
- Send the notes to apex-flash with the prompt below.
- Compile the rule with yara and fix syntax errors.
- Scan a known-clean corpus of your own binaries to find false positives.
- Scan your lab samples to confirm it matches what it should.
Authorisation boundary
This is for samples you are authorised to analyse, handled in an isolated lab. You are writing detection for them, not modifying or building malware, and the model should never be asked for working malicious code.
Sanitised sample input
Sample: 64-bit PE loader, about 180 KB, MZ header.
Strings seen: PDB path C:\build\agent\x64\Release\loader.pdb
Mutex: Global\UpdSvc_7f3a
UTF-16 string: "Software\Classes\CLSID\" (also appears in clean software)
User-Agent: Mozilla/5.0 (compatible; UpdSvc/2.1)
Not stable across builds: compile timestamp, a 16 byte keyThe prompt and the call
Per the YARA docs, a rule has optional meta and strings sections and a required condition. Text strings are case-sensitive ASCII unless you add nocase, wide or ascii, and conditions can test uint16(0) == 0x5A4D, filesize and 2 of ($a, $b).
Write one YARA rule from these analysis notes.
Use meta (author: TODO, description, date: TODO, reference: "internal analysis"), named strings with sensible modifiers (ascii, wide, nocase only where needed), and a condition that checks the MZ header, a filesize bound and at least two distinct indicators.
Do not use weak strings that appear in clean software. Mark any string you are unsure about with a comment.
After the rule, explain each condition clause and list likely false positives.
<notes>
...your notes...
</notes>import os
from openai import OpenAI
client = OpenAI(base_url="https://wildwestapi.com/v1",
api_key=os.environ["WILDWEST_API_KEY"])
resp = client.chat.completions.create(
model="apex-flash",
temperature=0.2,
messages=[
{"role": "system", "content": "You write precise YARA rules for defenders."},
{"role": "user", "content": open("yara_prompt.txt", encoding="utf-8").read()},
],
)
print(resp.choices[0].message.content)Keys look like sk-ww-...; keep yours in the WILDWEST_API_KEY environment variable, never in the script. Calls to /v1/chat/completions use the OpenAI format, billing is pay-as-you-go, and prompts are not retained on /v1.
What a good draft looks like
rule Loader_UpdSvc_Example
{
meta:
author = "TODO"
description = "Loader with UpdSvc mutex and PDB path"
strings:
$pdb = "C:\\build\\agent\\x64\\Release\\loader.pdb" ascii
$mutex = "Global\\UpdSvc_7f3a" ascii wide
$ua = "UpdSvc/2.1" ascii
condition:
uint16(0) == 0x5A4D and filesize < 500KB and 2 of them
}What to check
- The rule compiles. Then it must fire on the sample and stay silent on clean files; a rule that compiles is not a rule that detects.
- Backslashes in strings must be doubled in YARA. Check the model did this, because it is a common slip.
- The weak string (the CLSID path) should be dropped or demoted. If the draft leans on it, rewrite.
- Filesize bounds can hide variants. Widen them deliberately rather than guessing.
Both apex-flash and glm-5.3-flash-cyber are security-tuned models with a 1M-token context window, tool calling and vision. They are not uncensored models, and they are meant for defensive and authorised work like this.
Where this fits
Related work lives in malware analysis use cases. Pair rules with decompiled code explanations while you analyse, and Sigma rules for host behaviour. Tooling lists are on red team tools.
FAQ
Will it generate YARA rules from a sample file?
It reads text you send. For binaries, send your strings output and analysis notes, not the sample itself.
How do I measure false positives?
Scan a large set of clean binaries you own with yara and count matches. Do this before deploying any rule.
Why not ask it to write the malware behaviour too?
These models are security-tuned for defence and are not uncensored. The playbook only covers detection of samples you already hold.