Median Total Time
4.94s
Median TTFT
0.86s
Median Prefill TPS
74.77
Median Gen TPS
9.56
Context Size
262144
Quantization
r64
Engine
vllm
Creation Method
LoRA Finetune
Model Type
Gemma31B
Chat Template
Gemma4
Reasoning
Yes
Vision
Yes
Parameters
31B
Added At
7/25/2026
base_model:
This is a continued pretrain on the Gemma-4 31B base model. The resulting adapter was merged into google/gemma-4-31b-it. (Hence the "vanilla" part in the name.)
If this refuses, the non-vanilla model may be a better option.