Median Total Time
13.72s
Median TTFT
1.35s
Median Prefill TPS
91.12
Median Gen TPS
19.96
Context Size
262144
Quantization
r64 on FP8
Engine
vllm
Creation Method
Finetune
Model Type
Qwen35
Chat Template
Qwen3.5
Reasoning
Yes
Vision
Yes
Parameters
27B
Added At
7/20/2026
license: apache-2.0 datasets:
Blossom is a powerful open-source conversational large language model that provides reproducible post-training data, dedicated to delivering an open, powerful, and cost-effective locally accessible general-purpose model for everyone.
The Blossom-V6.4 series largely follows the V6.3 training recipe and uses the same training data, with a small number of multimodal samples added to preserve the multimodal capabilities of the Base models.
You can find the training data here: Blossom-V6.3-SFT-Stage1 (1 epoch)、Blossom-V6.3-SFT-Stage2 (3 epoch).
Primarily employs three cost-effective models: Deepseek-V3.1, Gemini 2.5 Flash, and Qwen3-235B-A22B-Instruct-2507 (denoted as A, B, C)—to regenerate responses under different scenarios using tailored synthesis strategies.
For example:
Additional rule-based filtering is applied, such as:
Further technical details will be released in the future. The data is synthesized by the 🌸BlossomData framework.
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "Azure99/Blossom-V6.4-27B"
model = AutoModelForCausalLM.from_pretrained(MODEL)
tokenizer = AutoTokenizer.from_pretrained(MODEL)
messages = [
{"role": "user", "content": "北京有什么好吃的"}
]
inputs = tokenizer.apply_chat_template(
messages,
return_tensors="pt",
return_dict=True,
).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=512)
generated_ids = generated_ids[:, inputs["input_ids"].shape[-1]:]
print(tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0])