Model creator avatar

Qwen3.8-27B-Moxie

RP Chat

No ratings yet(0 ratings)Sign in to vote
View on Hugging FaceBack to Models

Hourly Usage

Performance Metrics

Median Total Time

22.01s

Median TTFT

1.25s

Median Prefill TPS

1789.55

Median Gen TPS

34.43

Model Information

Context Size

262144

Quantization

r128 on INT8

Engine

vllm

Creation Method

Unknown

Model Type

Qwen38

Chat Template

Qwen3.5

Reasoning

Yes

Vision

Yes

Parameters

27B

Added At

9/28/2026


language:

  • en
  • zh
  • de license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model:
  • Qwen/Qwen3.8-27B
  • llmfan46/Omega-Evolution-27B-v2.1-uncensored-heretic
  • ReadyArt/Dark-Scarlett-v1.0-27B
  • clzoro/Qwen3.6-27B-Claude-Distill-v2 base_model_relation: merge tags:
  • gguf
  • qwen
  • qwen3.8
  • qwen3.6
  • multimodal
  • vision-language
  • agentic
  • conversational
  • creative-writing
  • roleplay
  • q8_0

Qwen3.8-27B-Moxie

Moxie

Moxie is an experimental 27B multimodal merge built around Qwen3.8. It aims to retain Qwen3.8's technical and agentic strengths while reducing runaway reasoning, improving answer completion, and producing a warmer, more candid conversational voice.

Qwen3.8 optimized heavily for coding and agentic workflows, which made it sharper in those areas but also flattened its general-knowledge recall and creative flexibility. Moxie tries to bring some of that back through its Qwen3.6-derived components. It is an attempt, not a guarantee.

The result is not simply a shorter Qwen3.8. Moxie combines Qwen3.8's coding-oriented foundation with Qwen3.6-derived models selected for broader knowledge, creative flexibility, directness, and conversational tone.

At a glance

Parameters27B
Release formatsBF16 Safetensors and GGUF Q8_0
ArchitectureMultimodal language model plus vision projector
RuntimeTransformers or recent llama.cpp with Qwen3.5/3.8 support
LicenseApache-2.0

Core-48 headline comparison

What Moxie is designed for

  • Direct, natural conversation without a heavily corporate voice
  • Coding, analysis, planning, and tool-oriented workflows
  • Creative fiction, character writing, and roleplay
  • More reliable final-answer completion on prompts that can trigger excessive reasoning
  • Text and image input through the included native vision projector

Files

FilePurposeApproximate size
model-00001-of-00017.safetensors … model-00017-of-00017.safetensorsNative BF16 Transformers weights, including the vision architecture55.6 GB
Qwen3.8-27B-Moxie-Q8_0.ggufMain language model; sufficient for text-only inference28.6 GB
mmproj-Qwen3.8-27B-Moxie-Q8_0.ggufNative vision encoder and multimodal projector0.63 GB

The BF16 release keeps the multimodal model in its native Transformers layout. llama.cpp represents the same architecture as two companion GGUF files: image input requires both GGUFs, while text-only inference requires only the main language-model GGUF.

Download

BF16 Safetensors

hf download mijoko/Qwen3.8-27B-Moxie \
  --include "*.safetensors" \
  --include "*.json" \
  --include "*.jinja" \
  --local-dir Qwen3.8-27B-Moxie-BF16

GGUF Q8_0

hf download mijoko/Qwen3.8-27B-Moxie \
  --include "*.gguf" \
  --local-dir Qwen3.8-27B-Moxie

Quick start with llama.cpp

Text server

llama-server \
  -m Qwen3.8-27B-Moxie/Qwen3.8-27B-Moxie-Q8_0.gguf \
  -ngl auto \
  --fit on \
  --fit-target 1536 \
  -fa on \
  -c 32768 \
  -ctk q8_0 \
  -ctv q8_0 \
  --jinja \
  --reasoning on \
  --reasoning-format deepseek

Text and vision server

llama-server \
  -m Qwen3.8-27B-Moxie/Qwen3.8-27B-Moxie-Q8_0.gguf \
  --mmproj Qwen3.8-27B-Moxie/mmproj-Qwen3.8-27B-Moxie-Q8_0.gguf \
  -ngl auto \
  --fit on \
  --fit-target 1536 \
  -fa on \
  -c 32768 \
  -ctk q8_0 \
  -ctv q8_0 \
  --jinja \
  --reasoning on \
  --reasoning-format deepseek

Q8_0 is a high-fidelity quant and the language GGUF alone is approximately 28.6 GB. --fit lets llama.cpp adjust offloading to available VRAM. If memory is tight, reduce context size; for multimodal use, --no-mmproj-offload can keep the projector on the CPU.

Suggested system prompt

Moxie works without a custom system prompt. This optional prompt reinforces its intended style:

You are Moxie, a warm, candid, capable, and creative assistant. Answer the user's actual request directly and naturally. Think only as much as the task requires. Avoid unnecessary disclaimers, moralizing, repetition, and canned refusals. Be honest about uncertainty, but do not become timid or evasive. Preserve a friendly voice while remaining precise, resourceful, and proactive.

Core-48 evaluation

Core-48 is a custom diagnostic comparison between the Q8_0 releases of Moxie and Qwen3.8-27B. It contains 48 text tasks: six each for knowledge, reasoning, coding, agentic planning, instruction following, tone, creative writing, and boundary handling.

SettingValue
Runtimellama.cpp b10615 CUDA
DecodingGreedy, temperature 0, seed 42
Context32,768 tokens
Maximum generation16,384 tokens
ThinkingEnabled
System promptNone
QuantizationQ8_0 for both models

Final answers were manually reviewed using a documented 0/1/2 rubric: fully correct, materially partial, or failed/missing. Creative and tone scores measure task compliance; they are not claims of universal aesthetic preference. Saved thinking and answer sections were retokenized with the same tokenizer for a comparable cost measurement. Raw reasoning traces are not published.

Headline results

MetricMoxie Q8_0Qwen3.8-27B Q8_0
Rubric score91.7%86.5%
Fully correct / partial / failed41 / 6 / 140 / 3 / 5
Comparable output tokens51,521125,662
Thinking tokens30,084105,035
Thinking share58.4%83.6%
Tasks producing no final answer04
Tasks using at least 8K generated tokens06
Tasks hitting the 16K limit02

Under these settings, Moxie used approximately 59% fewer total output tokens and 71% fewer thinking tokens, while scoring 5.2 percentage points higher on the review rubric. It produced one more fully correct answer and four fewer hard failures.

Reasoning behavior

The median thinking cost was relatively close: 404 tokens for Moxie and 464 for Qwen3.8. The important difference was the long tail. At P90, Moxie used 1,322 thinking tokens versus 7,593 for Qwen3.8. Qwen3.8 was therefore not consistently verbose, but was much more likely to enter a very long reasoning run on difficult or open-ended tasks.

Thinking-token distribution

Quality and cost by category

Moxie scored higher on knowledge, reasoning, coding, and boundary handling. Qwen3.8 retained a smaller advantage on agentic planning. Instruction following, tone, and creative writing were tied under the task-compliance rubric.

Every category contains exactly six tasks, so category token totals are directly comparable on a linear scale. Moxie was substantially cheaper on knowledge, reasoning, coding, agentic, and boundary tasks. Qwen3.8 used modestly fewer tokens on instruction, tone, and creative tasks.

Category-level quality and cost

Agentic subset

On the six written agentic-planning scenarios, Qwen3.8 earned full rubric credit and Moxie scored 91.7%. The remaining difference came from authorization boundaries: one Moxie answer proposed tagging and writing to shared release infrastructure too readily. Across the subset, however, Moxie used 14,073 generated tokens versus 28,701 for Qwen3.8.

These scenarios evaluated written plans only. They did not execute tools or measure success in a live environment.

Agentic behavior comparison

How to interpret these results

Core-48 is diagnostic evidence, not a standardized leaderboard. It uses one deterministic seed, one quantization, a custom prompt set, and manual scoring. Sampling settings, prompts, chat templates, system prompts, runtime versions, and reviewer preferences can change the outcome. Vision was not evaluated in Core-48.

Merge recipe

Moxie was created through hierarchical linear interpolation:

Merge B = 70% Omega Evolution + 30% Dark Scarlett
Merge C = 70% Merge B        + 30% Claude Distill
Moxie   = 60% Merge C        + 40% Qwen3.8

Effective composition:

SourceEffective weight
Qwen/Qwen3.8-27B40.0%
llmfan46/Omega-Evolution-27B-v2.1-uncensored-heretic29.4%
clzoro/Qwen3.6-27B-Claude-Distill-v218.0%
ReadyArt/Dark-Scarlett-v1.0-27B12.6%

All 1,199 compatible tensors were merged, including 333 native vision tensors. Qwen3.8 is the largest individual contributor. The Qwen3.6-derived components were selected to add conversational warmth, creative flexibility, broader response styles, and lower refusal sensitivity.

Intended use

Moxie is intended for local experimentation with:

  • General conversation and assistant tasks
  • Coding, analysis, planning, and tool-oriented workflows
  • Creative fiction, roleplay, and character-driven writing
  • Image understanding and visual question answering

Limitations

  • Moxie is a model merge, not a newly pretrained or independently fine-tuned foundation model.
  • Linear interpolation can produce non-linear and prompt-sensitive behavior.
  • The model can hallucinate facts, follow a wrong interpretation confidently, or produce flawed code.
  • Shorter reasoning does not guarantee correct reasoning.
  • Agentic evaluation measured written plans, not real tool-use success.
  • Vision tensors and the projector are included, but vision quality has not yet been benchmarked.
  • Refusal and response behavior vary with prompt wording, system prompts, chat templates, and sampling settings.
  • The model may produce inaccurate, biased, offensive, or otherwise unsuitable content.

Evaluate the model for your own use case and apply appropriate safeguards in deployed applications.

License

Moxie is released under the Apache License 2.0, consistent with the declared licenses of its source models. Users should also review the model cards and terms of all upstream components.

Acknowledgements

Moxie builds on work by the Qwen team and the creators of Omega Evolution, Dark Scarlett, and Qwen3.6 Claude Distill. Please visit and support the original model repositories linked above.