Median Total Time
28.90s
Median TTFT
6.56s
Median Prefill TPS
1058.80
Median Gen TPS
6.25
Context Size
262144
Quantization
r64 on INT8
Engine
vllm
Creation Method
Finetune
Model Type
Gemma31B
Chat Template
Gemma4
Reasoning
Yes
Vision
Yes
Parameters
31B
Added At
9/26/2026
license: apache-2.0 datasets:
The high-fidelity 31B Aura release for Transformers inference, evaluation, and research.
Aura Large v1.1 BF16 is the full-precision merged edition of Aura Large v1.1. It is derived from Google’s Gemma 4 31B IT checkpoint and combines the Aura personality and ablation adapters into one model.
The v1.1 release has seen significant improvements to the personality adapter, merged and applied with the same ablation adapter as v1.
This release is intended for Transformers-based inference, evaluation, further research, and deployment on GPUs with sufficient memory.
Aura is designed for local and on-device deployment across multiple tasks. Aura can serve as a companion or friend, as deemed appropriate by the user, while retaining the broader capabilities of Gemma 4 for:
Aura Large v1.1 occupies the largest position in the Aura family, providing greater capability than the Medium release while remaining practical for high-end consumer hardware.