Median Total Time
13.22s
Median TTFT
7.52s
Median Prefill TPS
991.68
Median Gen TPS
24.05
Context Size
262144
Quantization
r64 on INT8
Engine
vllm
Creation Method
Finetune
Model Type
Gemma31B
Chat Template
Gemma4
Reasoning
Yes
Vision
Yes
Parameters
31B
Added At
7/25/2026
base_model:
This is a continued pretrain on the Gemma-4 31B base model. The resulting adapter was merged into google/gemma-4-31b-it. (Hence the "vanilla" part in the name.)
If this refuses, the non-vanilla model may be a better option.