Is there an uncensored NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)?
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) from NVIDIA is ranked 75 of 100 distinct active open-weight models in this Artificial Analysis snapshot.
Research and availability status
This is a research guide, not an available Wild West API model listing. No uncensored or abliterated NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) endpoint is confirmed here. An open-weight entry on AA does not establish that a derivative exists, works with a particular harness, or is hosted by Wild West API.
“Uncensored” describes reduced refusal behaviour or a service’s policy, not a standardized editing method. “Abliterated” refers to an intervention intended to reduce refusal-related behaviour through model-internal edits.
What rank 75 actually tells you
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) scores above 25 of the 100 selected entries: a strictly-lower-score percentile of 25% in this snapshot. No other selected entry shares its exact index score. This entry is not flagged as an estimated index in the source data.
It falls in the lower half of this top-100 selection. That does not make it unusable: a narrowly defined task can differ from the aggregate evaluation. However, this data alone supplies no efficiency, hardware or price reason to choose it over a stronger entry.
Base-model benchmark evidence
These are measurements attached to the evaluated base model, not an uncensored derivative. Missing values mean not reported here, never a zero score. Same-benchmark comparisons below use only entries with a reported measurement.
| Benchmark | Reported result | Within-snapshot comparison | Task coverage |
|---|---|---|---|
| GPQA | 75.7% | 51 of 96 measured entries score lower on this same benchmark | graduate-level science questions |
| Humanity’s Last Exam | 11.4% | 50 of 100 measured entries score lower on this same benchmark | a broad set of difficult expert questions |
| SciCode | 30.6% | 9 of 56 measured entries score lower on this same benchmark | scientific coding tasks |
| IFBench | 71.1% | 56 of 65 measured entries score lower on this same benchmark | instruction-following constraints |
| τ²-Bench | 40.9% | 31 of 64 measured entries score lower on this same benchmark | interactive agent tasks |
| Terminal-Bench Hard | 13.6% | 37 of 63 measured entries score lower on this same benchmark | hard terminal-based tasks |
The highest numerical reported percentage for NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) is GPQA at 75.7%; the lowest is Humanity’s Last Exam at 11.4%. These are numerical extremes, not evidence that one skill is stronger than another: benchmark percentages cannot be compared as if the tests had equal difficulty, scoring rules or task coverage. Compare GPQA with that same benchmark on another model, not with a different test.
Context window and a practical evaluation workload
At 1,000,000 tokens, a candidate workload is a large document collection or a repository with supporting specifications. A useful test would ask the model to trace a claim across distant files, then verify every citation. A large advertised window does not establish accurate retrieval throughout that window, and packing it fully can change latency and cost.
81 of the 100 entries have a smaller reported window than NVIDIA Nemotron 3 Nano 30B A3B (Reasoning). As a planning example, allocating 75% of its reported window to source material gives 750,000 tokens, leaving 250,000 for instructions, conversation, tool messages and output combined. This is arithmetic for a test budget, not an endpoint limit or an assertion that output can use the entire remainder.
Reasoning settings and release comparisons
The 2.05-point spread is a setting comparison, not an abliteration effect. A reasoning-budget label is not a separate base-family entry in the count of 100.
| Evaluated setting | Intelligence Index | Reasoning setting |
|---|---|---|
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) | 8.90 | Yes |
| NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) | 6.85 | No |
Other entries credited to NVIDIA provide a publisher-level comparison, not proof of a shared architecture or training recipe. Nemotron 3 Nano Omni 30B A3B Reasoning ranks 63, scores 10.25 and reports 256,000 context tokens; Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) ranks 89, scores 7.53 and reports 128,000 context tokens; NVIDIA Nemotron Nano 12B v2 VL (Reasoning) ranks 92, scores 7.48 and reports 128,000 context tokens. Compare exact releases and evaluation settings rather than carrying a family’s reputation over to this model.
Nearby ranked alternatives
The closest ranks give a bounded shortlist for evaluating NVIDIA Nemotron 3 Nano 30B A3B (Reasoning). A small index gap is not a statistical significance claim; compare task outcomes and deployment conditions directly.
| Alternative | Rank | Index and gap from this model | Context |
|---|---|---|---|
| Tri-21B-Think | 73 | 8.99 (+0.09 points) | 32,000 tokens |
| Gemma 4 E4B (Reasoning) | 74 | 8.91 (+0.02 points) | 128,000 tokens |
| MiniCPM5-1B (Reasoning) | 76 | 8.80 (-0.10 points) | 128,000 tokens |
| Sarvam 105B (High) | 77 | 8.79 (-0.11 points) | 128,000 tokens |
How to evaluate an uncensored or abliterated candidate
For NVIDIA Nemotron 3 Nano 30B A3B (Reasoning), first locate the exact base release and any claimed derivative’s model card. Verify access, redistribution and usage terms at their original sources: AA’s open-weight classification does not assert an open-source license. Record the base revision, derivative revision, runtime and reasoning setting so your comparison can be reproduced.
Weight-based abliteration generally studies differences between activations on refusal-inducing and ordinary prompts, identifies candidate refusal-related directions, and intervenes on model internals. It is not equivalent to changing a system prompt or removing an API filter. Whether a particular intervention is supported depends on the model and implementation; this guide supplies no architecture-specific recipe or claim that the method has been validated on NVIDIA Nemotron 3 Nano 30B A3B (Reasoning).
Run the base and candidate on the same permitted task set before and after changes. Measure refusal behaviour separately from factual errors, instruction following, coding correctness and tool execution. Use GPQA as one reference for task coverage, while keeping its published percentage separate from your own results. Track any regression and repeat with the exact context length and setting your application needs. A lower refusal rate is not evidence that all other behaviour survived unchanged.
Sources and scope
Snapshot retrieved 2026-10-06T06:17:07.812043+00:00. AA’s NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) model entry supplies the evaluated setting and measurements; the Artificial Analysis leaderboard is the population source. Top 100 distinct active open-weight model names by Artificial Analysis Intelligence Index. Reasoning-budget parentheticals are removed for grouping; the strongest scored setting represents each group. Open weights is AA's classification, not an open-source-license assertion. Scores describe the evaluated base models, not uncensored derivatives. Rank is within this selected active open-weight population, not the complete leaderboard.
For the research behind refusal-direction interventions, see Arditi et al., Refusal in Language Models Is Mediated by a Single Direction. This general research source does not establish successful abliteration of this release. Continue with the abliteration explanation or the current Wild West API line; only the latter lists products for sale.
Questions about NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)
Is NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) abliterated on Wild West API?
Not confirmed by this guide. This page documents an AA open-weight research entry, not a hosted derivative. Check the current model line for available products.
Does the 8.90 index score apply to an uncensored derivative?
No. It describes AA’s evaluated NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) base-model setting. Changes to weights, prompts, quantization, tools or reasoning budgets require fresh evaluation before carrying the score over.
Can I use the full 1,000,000-token window?
That is the context figure in the AA snapshot, not a guarantee for every deployment. Confirm the actual endpoint or runtime limits and reserve space for output and tool messages. Test retrieval accuracy near the ends of long inputs.