Model creator avatar

Qwen3.8-27B-Kiwen1.1

All-rounder

No ratings yet(0 ratings)Sign in to vote
View on Hugging FaceBack to Models

Hourly Usage

Performance Metrics

Median Total Time

22.01s

Median TTFT

1.25s

Median Prefill TPS

1789.55

Median Gen TPS

34.43

Model Information

Context Size

262144

Quantization

r128 on INT8

Engine

vllm

Creation Method

Unknown

Model Type

Qwen38

Chat Template

Qwen3.5

Reasoning

Yes

Vision

Yes

Parameters

27B

Added At

9/28/2026


license: apache-2.0 base_model: Qwen/Qwen3.8-27B pipeline_tag: image-text-to-text library_name: transformers tags:

  • qwen3.8
  • reasoning
  • agentic
  • vietnamese
  • math
  • coding
  • kimi
  • distillation
  • kiwen1.1
  • transformer

Kiwen1.1-27B

Kiwen1.1 is a fine-tuned version of Qwen3.8-27B, trained on long chain-of-thought reasoning traces from Kimi K3 — with a particular focus on reasoning, coding, tool use, and instruction following. But the goal wasn't simply to make the model think longer. The goal was to make it think better.

This model isn't an attempt to build the biggest model.

It's an attempt to make a 27-billion-parameter model think harder, follow instructions better, and act more reliably.

News

  • Traing focus more in coding task
  • Improve instruct response of base model
  • agent workflow internal tasks

Results

lm-evaluation-harness 0.4.12, chat template applied, max_gen_toks=4096, full datasets. Base measured under the identical harness.

Qwen3.8-27BKiwen-27BKiwen1.1-27B
GSM8K strict67.472.7896.4
GSM8K flexible74.984.3896.7
IFEval prompt strict80.484.2983.9
IFEval inst strict82.586.4587.5
IFEval prompt loose83.286.6987.2
IFEval inst loose84.388.0189.7
VMLU val (744)83.586.0284.8

GSM8K moves by 29 points. That is the headline number and it is real, on the full 1,319-item set. Base and fine-tune were run back to back in the same job so the two columns share their conditions. A second independent run of the fine-tune scored 96.2 / 96.4, which puts the run-to-run spread around 0.3.

VMLU moves much less: +1.3 over the base, against Kiwen-27B's +2.5. The gain is real but small. The training mix is weighted toward reasoning and tool use, not toward Vietnamese factual recall, so a model tuned directly for that recall stays ahead. Base VMLU breaks down as STEM 93.4, Social Science 82.4, Other 78.6, Humanity 74.7, and the humanities gap is where the headroom is.

Internal benchmark

On an internal 20-task suite the model won 5, lost 6 and tied 5, median delta 0.0. The wins are concentrated in classification and routing:

TaskDelta
Translation+22.9
Intent classification+20.9
Intent routing+20.7
Multi-turn tool calling+15.1

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

m = "beyoru/Kiwen1.1-27B"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, dtype="auto", device_map="auto")

msgs = [{"role": "user", "content": "Natalia sold clips to 48 friends in April, "
                                     "and half as many in May. How many total?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True,
                              enable_thinking=True, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(ids, max_new_tokens=4096)[0][ids.shape[-1]:]))

Set enable_thinking=False for extraction, classification and formatting tasks. The model was trained with both modes and respects the flag.

Serving with SGLang:

python -m sglang.launch_server --model-path beyoru/Kiwen1.1-27B \
  --context-length 262144 

Notes:

  • Found some issues in response quality when using with dflash2 - make sure you don't use this model with it

Citation

@misc{kiwen27bk3,
  title  = {Kiwen1.1-27B},
  author = {beyoru},
  year   = {2026},
  url    = {https://huggingface.co/beyoru/Kiwen1.1-27B}
}

License & Attribution

Built on:

Kiwen1.1-27B is released under the Apache-2.0 license.