Median Total Time
23.56s
Median TTFT
3.30s
Median Prefill TPS
2467.25
Median Gen TPS
24.47
Context Size
262144
Quantization
r128 on INT8
Engine
vllm
Creation Method
Unknown
Model Type
Qwen38B
Chat Template
Qwen3.5
Reasoning
Yes
Vision
Yes
Parameters
27B
Added At
10/1/2026
license: apache-2.0 base_model: Qwen/Qwen3.8-27B tags:
grug think small, answer big.
normal model think like this:
Okay, so the user wants me to carefully consider the best approach here. Let me think about this step by step. First, I should consider what data structure...
grug think like this:
Sort numbers once; adjacent differences sufficient. Any diff < threshold -> True; else False. Edge len < 2 -> False.
same reasoning. same steps. same answer quality. grug throw grammar padding in fire, keep brain meat. answer come out normal english — only inside voice is grug.
new rock under grug: Qwen3.8-27B. old grug sit on Qwen3.6.


base model burn 559 token thinking about HumanEval. grug burn 79.5. same kind of answer. on agent step base burn 108.5, grug burn 20.

this is the big one. base model can call a tool — it call a valid tool 98.5% of time. but base pick the right tool only 23.5% of time. grug pick right tool 97.1%.
medium reasoning effort, full benchmark sets (HumanEval 164, MBPP 100, GSM8K 200, MATH-500 150, agentic 68, recovery 80, repetition 43). same harness, same settings, every column.
| benchmark | Qwen3.8 base | grug v1 | grug v1.1 |
|---|---|---|---|
| HumanEval | 98.2 | 87.8 | 94.5 |
| MBPP | 93.0 | 84.0 | 88.0 |
| GSM8K | 95.5 | 96.5 | 92.5 |
| MATH-500 | 78.0 | 64.7 | 72.7 |
| repetition stress | 76.7 | 81.4 | 88.4 |
| agentic — valid call | 98.5 | 100.0 | 100.0 |
| agentic — right tool | 23.5 | 95.6 | 97.1 |
| agentic — args valid | 98.5 | 100.0 | 100.0 |
| recovery — valid call | 100.0 | 100.0 | 100.0 |
| recovery — right tool | 32.5 | 90.0 | 82.5 |
| loops / unclosed think | — | — | 0 / 0 |
mean reasoning token per answer:
| benchmark | Qwen3.8 base | grug v1 | grug v1.1 |
|---|---|---|---|
| HumanEval | 559.0 | 42.2 | 79.5 |
| MBPP | 656.4 | 34.6 | 301.2 |
| GSM8K | 204.5 | 76.0 | 64.9 |
| MATH-500 | 750.5 | 163.2 | 190.3 |
| agentic step | 108.5 | 29.4 | 20.0 |
| failure recovery | 78.2 | 33.4 | 24.0 |

Qwen3.8 template take reasoning_effort. same weight, three setting. grug measure
all three on this exact released model, full sets:
| probe | low | medium | xhigh |
|---|---|---|---|
| HumanEval | 87.2 | 94.5 | 92.1 |
| MBPP | 85.0 | 88.0 | 81.0 |
| GSM8K | 91.5 | 92.5 | 90.5 |
| MATH-500 | 70.0 | 72.7 | 72.0 |
| agentic — right tool | 89.7 | 97.1 | 76.5 |
| recovery — right tool | 85.0 | 82.5 | 86.2 |
| repetition stress | 93.0 | 88.4 | 90.7 |
| summed think token | 464 | 680 | 628 |
more effort is NOT more better. medium win 5 of 7 probe. xhigh actively hurt tool pick — 76.5 vs 97.1, drop 20 point, because xhigh instruction tell model "think carefully, consider alternative", and that push agent toward writing essay instead of calling tool. base model also get worse at xhigh (HumanEval 98.2 -> 92.7). so grug say: use medium.
low is real budget option though: 464 think token total instead of 680 (-32%),
and still 85.0 MBPP / 93.0 repetition. pay ~7 point HumanEval for it.
(old grug v1 cannot do this at all — Qwen3.6 template ignore reasoning_effort,
all three setting render byte-identical prompt. dial is new.)
grug not hide bruise.
because gain over v1 is honest-small. new base rock, big win on tool choice and repetition and MATH-500, but GSM8K step back. that is a point-one, not a two. grug not put big number on small step.
step 3 matter more than it sound. at full strength the adapter overshoot — it push so hard toward tool behaviour that it break code:
| adapter strength | HumanEval | right tool |
|---|---|---|
| 1.0x | 84.8 | 88.2 |
| 0.7x | 90.9 | 94.1 |
| 0.5x (shipped) | 94.5 | 97.1 |
less adapter = better code AND better tool pick, both at once. grug learn: more push not always more better.
first build of this model score fine on code and then collapse on tool call — 20.6% valid call. three rot in training data, none in model:
<think> block from
reasoning_content field. old data hid think inside content as <think> tag.
so every row train as <think>\n\n</think> then second literal <think> —
empty think, then duplicate. 59% of supervised turn carried empty think. that
why model sometimes shut reasoning instantly then reason inside answer.old 35b brother repeat word until cave fall down. grug train on BOTH world: think-in-history trajectories AND stripped-history variants (old think gone, exactly like real agent framework replay). then repetition stress gauntlet before release: greedy long-form, deep think-stripped agent replay, multi-turn continuation. zero loop, 100% think closed.
from transformers import AutoTokenizer, AutoModelForImageTextToText
tok = AutoTokenizer.from_pretrained("ProCreations/grug-v1.1-qwen-3.8-27b")
model = AutoModelForImageTextToText.from_pretrained(
"ProCreations/grug-v1.1-qwen-3.8-27b", dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "write a function that flattens a nested list"}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
reasoning_effort="medium", return_tensors="pt")
print(tok.decode(model.generate(ids.to(model.device), max_new_tokens=512)[0]))
use reasoning_effort="medium". every number on this card is medium. grug
tuned there. low and xhigh work but are not what grug measured for release.
GGUF: ProCreations/grug-v1.1-qwen-3.8-27b-gguf
(with mmproj for vision).
<function=> shape, same as grug v1.apache-2.0, like the rock it stand on.