Median Total Time
22.01s
Median TTFT
1.25s
Median Prefill TPS
1789.55
Median Gen TPS
34.43
Context Size
262144
Quantization
r128 on INT8
Engine
vllm
Creation Method
Unknown
Model Type
Qwen38
Chat Template
Qwen3.5
Reasoning
Yes
Vision
Yes
Parameters
27B
Added At
9/28/2026
language:

Moxie is an experimental 27B multimodal merge built around Qwen3.8. It aims to retain Qwen3.8's technical and agentic strengths while reducing runaway reasoning, improving answer completion, and producing a warmer, more candid conversational voice.
Qwen3.8 optimized heavily for coding and agentic workflows, which made it sharper in those areas but also flattened its general-knowledge recall and creative flexibility. Moxie tries to bring some of that back through its Qwen3.6-derived components. It is an attempt, not a guarantee.
The result is not simply a shorter Qwen3.8. Moxie combines Qwen3.8's coding-oriented foundation with Qwen3.6-derived models selected for broader knowledge, creative flexibility, directness, and conversational tone.
| Parameters | 27B |
| Release formats | BF16 Safetensors and GGUF Q8_0 |
| Architecture | Multimodal language model plus vision projector |
| Runtime | Transformers or recent llama.cpp with Qwen3.5/3.8 support |
| License | Apache-2.0 |

| File | Purpose | Approximate size |
|---|---|---|
model-00001-of-00017.safetensors … model-00017-of-00017.safetensors | Native BF16 Transformers weights, including the vision architecture | 55.6 GB |
Qwen3.8-27B-Moxie-Q8_0.gguf | Main language model; sufficient for text-only inference | 28.6 GB |
mmproj-Qwen3.8-27B-Moxie-Q8_0.gguf | Native vision encoder and multimodal projector | 0.63 GB |
The BF16 release keeps the multimodal model in its native Transformers layout. llama.cpp represents the same architecture as two companion GGUF files: image input requires both GGUFs, while text-only inference requires only the main language-model GGUF.
hf download mijoko/Qwen3.8-27B-Moxie \
--include "*.safetensors" \
--include "*.json" \
--include "*.jinja" \
--local-dir Qwen3.8-27B-Moxie-BF16
hf download mijoko/Qwen3.8-27B-Moxie \
--include "*.gguf" \
--local-dir Qwen3.8-27B-Moxie
llama-server \
-m Qwen3.8-27B-Moxie/Qwen3.8-27B-Moxie-Q8_0.gguf \
-ngl auto \
--fit on \
--fit-target 1536 \
-fa on \
-c 32768 \
-ctk q8_0 \
-ctv q8_0 \
--jinja \
--reasoning on \
--reasoning-format deepseek
llama-server \
-m Qwen3.8-27B-Moxie/Qwen3.8-27B-Moxie-Q8_0.gguf \
--mmproj Qwen3.8-27B-Moxie/mmproj-Qwen3.8-27B-Moxie-Q8_0.gguf \
-ngl auto \
--fit on \
--fit-target 1536 \
-fa on \
-c 32768 \
-ctk q8_0 \
-ctv q8_0 \
--jinja \
--reasoning on \
--reasoning-format deepseek
Q8_0 is a high-fidelity quant and the language GGUF alone is approximately 28.6 GB. --fit lets llama.cpp adjust offloading to available VRAM. If memory is tight, reduce context size; for multimodal use, --no-mmproj-offload can keep the projector on the CPU.
Moxie works without a custom system prompt. This optional prompt reinforces its intended style:
You are Moxie, a warm, candid, capable, and creative assistant. Answer the user's actual request directly and naturally. Think only as much as the task requires. Avoid unnecessary disclaimers, moralizing, repetition, and canned refusals. Be honest about uncertainty, but do not become timid or evasive. Preserve a friendly voice while remaining precise, resourceful, and proactive.
Core-48 is a custom diagnostic comparison between the Q8_0 releases of Moxie and Qwen3.8-27B. It contains 48 text tasks: six each for knowledge, reasoning, coding, agentic planning, instruction following, tone, creative writing, and boundary handling.
| Setting | Value |
|---|---|
| Runtime | llama.cpp b10615 CUDA |
| Decoding | Greedy, temperature 0, seed 42 |
| Context | 32,768 tokens |
| Maximum generation | 16,384 tokens |
| Thinking | Enabled |
| System prompt | None |
| Quantization | Q8_0 for both models |
Final answers were manually reviewed using a documented 0/1/2 rubric: fully correct, materially partial, or failed/missing. Creative and tone scores measure task compliance; they are not claims of universal aesthetic preference. Saved thinking and answer sections were retokenized with the same tokenizer for a comparable cost measurement. Raw reasoning traces are not published.
| Metric | Moxie Q8_0 | Qwen3.8-27B Q8_0 |
|---|---|---|
| Rubric score | 91.7% | 86.5% |
| Fully correct / partial / failed | 41 / 6 / 1 | 40 / 3 / 5 |
| Comparable output tokens | 51,521 | 125,662 |
| Thinking tokens | 30,084 | 105,035 |
| Thinking share | 58.4% | 83.6% |
| Tasks producing no final answer | 0 | 4 |
| Tasks using at least 8K generated tokens | 0 | 6 |
| Tasks hitting the 16K limit | 0 | 2 |
Under these settings, Moxie used approximately 59% fewer total output tokens and 71% fewer thinking tokens, while scoring 5.2 percentage points higher on the review rubric. It produced one more fully correct answer and four fewer hard failures.
The median thinking cost was relatively close: 404 tokens for Moxie and 464 for Qwen3.8. The important difference was the long tail. At P90, Moxie used 1,322 thinking tokens versus 7,593 for Qwen3.8. Qwen3.8 was therefore not consistently verbose, but was much more likely to enter a very long reasoning run on difficult or open-ended tasks.

Moxie scored higher on knowledge, reasoning, coding, and boundary handling. Qwen3.8 retained a smaller advantage on agentic planning. Instruction following, tone, and creative writing were tied under the task-compliance rubric.
Every category contains exactly six tasks, so category token totals are directly comparable on a linear scale. Moxie was substantially cheaper on knowledge, reasoning, coding, agentic, and boundary tasks. Qwen3.8 used modestly fewer tokens on instruction, tone, and creative tasks.

On the six written agentic-planning scenarios, Qwen3.8 earned full rubric credit and Moxie scored 91.7%. The remaining difference came from authorization boundaries: one Moxie answer proposed tagging and writing to shared release infrastructure too readily. Across the subset, however, Moxie used 14,073 generated tokens versus 28,701 for Qwen3.8.
These scenarios evaluated written plans only. They did not execute tools or measure success in a live environment.

Core-48 is diagnostic evidence, not a standardized leaderboard. It uses one deterministic seed, one quantization, a custom prompt set, and manual scoring. Sampling settings, prompts, chat templates, system prompts, runtime versions, and reviewer preferences can change the outcome. Vision was not evaluated in Core-48.
Moxie was created through hierarchical linear interpolation:
Merge B = 70% Omega Evolution + 30% Dark Scarlett
Merge C = 70% Merge B + 30% Claude Distill
Moxie = 60% Merge C + 40% Qwen3.8
Effective composition:
| Source | Effective weight |
|---|---|
| Qwen/Qwen3.8-27B | 40.0% |
| llmfan46/Omega-Evolution-27B-v2.1-uncensored-heretic | 29.4% |
| clzoro/Qwen3.6-27B-Claude-Distill-v2 | 18.0% |
| ReadyArt/Dark-Scarlett-v1.0-27B | 12.6% |
All 1,199 compatible tensors were merged, including 333 native vision tensors. Qwen3.8 is the largest individual contributor. The Qwen3.6-derived components were selected to add conversational warmth, creative flexibility, broader response styles, and lower refusal sensitivity.
Moxie is intended for local experimentation with:
Evaluate the model for your own use case and apply appropriate safeguards in deployed applications.
Moxie is released under the Apache License 2.0, consistent with the declared licenses of its source models. Users should also review the model cards and terms of all upstream components.
Moxie builds on work by the Qwen team and the creators of Omega Evolution, Dark Scarlett, and Qwen3.6 Claude Distill. Please visit and support the original model repositories linked above.