Wild West API

Is there an uncensored Qwen3.8 27B (Xhigh)?

Qwen3.8 27B (Xhigh) from Alibaba is ranked 10 of 100 distinct active open-weight models in this Artificial Analysis snapshot.

Research and availability status

The current line contains Qwen3.8 27B Xploded, an offered Xploded variant associated with this exact release. It is a separate product from AA’s evaluated Qwen3.8 27B (Xhigh) setting. The measurements below must not be treated as its benchmark results or a claim that Wild West API made the weight edits.

“Uncensored” describes reduced refusal behaviour or a service’s policy, not a standardized editing method. “Abliterated” refers to an intervention intended to reduce refusal-related behaviour through model-internal edits.

What rank 10 actually tells you

Qwen3.8 27B (Xhigh) scores above 90 of the 100 selected entries: a strictly-lower-score percentile of 90% in this snapshot. No other selected entry shares its exact index score. This entry is not flagged as an estimated index in the source data.

Its rank puts it in the first ten entries of this selected population. Start comparisons with the nearest high-scoring models rather than assuming the lead transfers to every task; the index is an aggregate, not a refusal-rate, reliability or throughput measurement.

Base-model benchmark evidence

These are measurements attached to the evaluated base model, not an uncensored derivative. Missing values mean not reported here, never a zero score. Same-benchmark comparisons below use only entries with a reported measurement.

BenchmarkReported resultWithin-snapshot comparisonTask coverage
GPQA90.5%88 of 96 measured entries score lower on this same benchmarkgraduate-level science questions
Humanity’s Last Exam33.9%87 of 100 measured entries score lower on this same benchmarka broad set of difficult expert questions
SciCode46.6%40 of 56 measured entries score lower on this same benchmarkscientific coding tasks
IFBenchNot reportedNo comparison availableinstruction-following constraints
τ²-BenchNot reportedNo comparison availableinteractive agent tasks
Terminal-Bench HardNot reportedNo comparison availablehard terminal-based tasks

The highest numerical reported percentage for Qwen3.8 27B (Xhigh) is GPQA at 90.5%; the lowest is Humanity’s Last Exam at 33.9%. These are numerical extremes, not evidence that one skill is stronger than another: benchmark percentages cannot be compared as if the tests had equal difficulty, scoring rules or task coverage. Compare GPQA with that same benchmark on another model, not with a different test.

Context window and a practical evaluation workload

The 256,000-token window makes multi-document review or a focused repository investigation worth testing. Reserve room for instructions, tool responses and the answer before loading source material. Compare a full-context run with a retrieval-based run on the same questions; the window alone cannot tell you which produces better grounded answers.

36 of the 100 entries have a smaller reported window than Qwen3.8 27B (Xhigh). As a planning example, allocating 75% of its reported window to source material gives 192,000 tokens, leaving 64,000 for instructions, conversation, tool messages and output combined. This is arithmetic for a test budget, not an endpoint limit or an assertion that output can use the entire remainder.

Reasoning settings and release comparisons

The 13.55-point spread is a setting comparison, not an abliteration effect. A reasoning-budget label is not a separate base-family entry in the count of 100.

Evaluated settingIntelligence IndexReasoning setting
Qwen3.8 27B (Xhigh)33.70Yes
Qwen3.8 27B (Medium)27.55Yes
Qwen3.8 27B (Low)26.20Yes
Qwen3.8 27B (Non-reasoning)20.15No

Other entries credited to Alibaba provide a publisher-level comparison, not proof of a shared architecture or training recipe. Qwen3.8-Flash-Next ranks 6, scores 39.82 and reports 256,000 context tokens; Qwen3.8 2.4T A95B ranks 5, scores 39.89 and reports 262,000 context tokens; Qwen3.5 397B A17B (Non-reasoning) ranks 26, scores 21.44 and reports 262,144 context tokens. Compare exact releases and evaluation settings rather than carrying a family’s reputation over to this model.

Nearby ranked alternatives

The closest ranks give a bounded shortlist for evaluating Qwen3.8 27B (Xhigh). A small index gap is not a statistical significance claim; compare task outcomes and deployment conditions directly.

AlternativeRankIndex and gap from this modelContext
MiMo-V2.6-Flash837.88 (+4.19 points)1,000,000 tokens
DeepSeek V4 Pro 0813 (Max)936.00 (+2.30 points)1,000,000 tokens
Motif 31133.57 (-0.13 points)262,144 tokens
K2 Horizon 375B A23B1230.50 (-3.19 points)524,288 tokens

How to evaluate an uncensored or abliterated candidate

For Qwen3.8 27B (Xhigh), first locate the exact base release and any claimed derivative’s model card. Verify access, redistribution and usage terms at their original sources: AA’s open-weight classification does not assert an open-source license. Record the base revision, derivative revision, runtime and reasoning setting so your comparison can be reproduced.

Weight-based abliteration generally studies differences between activations on refusal-inducing and ordinary prompts, identifies candidate refusal-related directions, and intervenes on model internals. It is not equivalent to changing a system prompt or removing an API filter. Whether a particular intervention is supported depends on the model and implementation; this guide supplies no architecture-specific recipe or claim that the method has been validated on Qwen3.8 27B (Xhigh).

Run the base and candidate on the same permitted task set before and after changes. Measure refusal behaviour separately from factual errors, instruction following, coding correctness and tool execution. Use GPQA as one reference for task coverage, while keeping its published percentage separate from your own results. Track any regression and repeat with the exact context length and setting your application needs. A lower refusal rate is not evidence that all other behaviour survived unchanged.

Sources and scope

Snapshot retrieved 2026-10-06T06:17:07.812043+00:00. AA’s Qwen3.8 27B (Xhigh) model entry supplies the evaluated setting and measurements; the Artificial Analysis leaderboard is the population source. Top 100 distinct active open-weight model names by Artificial Analysis Intelligence Index. Reasoning-budget parentheticals are removed for grouping; the strongest scored setting represents each group. Open weights is AA's classification, not an open-source-license assertion. Scores describe the evaluated base models, not uncensored derivatives. Rank is within this selected active open-weight population, not the complete leaderboard.

For the research behind refusal-direction interventions, see Arditi et al., Refusal in Language Models Is Mediated by a Single Direction. This general research source does not establish successful abliteration of this release. Continue with the abliteration explanation or the current Wild West API line; only the latter lists products for sale.

Questions about Qwen3.8 27B (Xhigh)

Is Qwen3.8 27B (Xhigh) abliterated on Wild West API?

An associated Xploded variant is present on the current line, but this guide does not establish its editing method. Check its product page and the current /v1/models listing. AA’s base-model score is not a derivative score.

Does the 33.70 index score apply to an uncensored derivative?

No. It describes AA’s evaluated Qwen3.8 27B (Xhigh) base-model setting. Changes to weights, prompts, quantization, tools or reasoning budgets require fresh evaluation before carrying the score over.

Can I use the full 256,000-token window?

That is the context figure in the AA snapshot, not a guarantee for every deployment. Confirm the actual endpoint or runtime limits and reserve space for output and tool messages. Test retrieval accuracy near the ends of long inputs.

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.