Median Total Time
13.42s
Median TTFT
7.55s
Median Prefill TPS
988.03
Median Gen TPS
23.85
Context Size
262144
Quantization
r64 on INT8
Engine
vllm
Creation Method
Finetune
Model Type
Gemma31B
Chat Template
Gemma4
Reasoning
Yes
Vision
Yes
Parameters
31B
Added At
7/19/2026
license: apache-2.0 language:
google/gemma-4-31B-it finetuned for literary, novelistic prose — a Gemma 4 entry in the Gutenberg series.
This pushes the (already strong) base toward a literary-fiction register: story and interiority over static description, controlled pacing over relentless adjective-stacking, and an active dispreference for "AI slop" phrasing.
ORPO (Odds Ratio Preference Optimization) on the full
schneewolflabs/Athanorlite-DPO
(14,816 preference pairs) — a superset of the Gutenberg "Encore" recipe that
bundles jondurbin/gutenberg-dpo-v0.1, nbeerbower/gutenberg2-dpo,
gutenberg-moderne-dpo, human-writing-dpo, synthetic-fiction-dpo,
Arkhaios-DPO, Purpura-DPO, Schule-DPO,
sam-paech/gutenberg3, plus truthy / physical-reasoning / theory-of-mind
balance sets.
| Method | ORPO, β = 0.1 |
| Adapter | LoRA r=64 (text decoder only), merged to full |
| LR | 5e-5, cosine, 0.05 warmup |
| Epochs | 1 |
| Effective batch | 32 |
| Max length | 2048 |
| Optimizer | paged_adamw_8bit, bf16, grad-checkpointing |
| Hardware | 1× NVIDIA GB10 (DGX Spark, 128 GB unified) |
| Trainer | Merlina (grimoire ORPO) |
Training trajectory (clean convergence over ~3.5 days, 454 steps):
| start | end | |
|---|---|---|
| eval/loss | 2.4445 | 2.2052 |
| reward_accuracy | 0.175 | 0.9125 |
| reward_margin | −0.076 | +0.455 |
The reward_accuracy arc (0.18 → 0.50 → 0.91) reflects the model learning to
prefer the literary chosen text while actively suppressing the rejected
slop — the intended Gutenberg dynamic.
Apache-2.0 (matching the Gemma 4 base). Constituent training datasets carry their own licenses (see the Athanorlite-DPO card).