Gemma-4-31B-Lilith-v1.0

Creative model

View on Hugging FaceBack to Models

Hourly Usage

Performance Metrics

Median Total Time

24.68s

Median TTFT

6.51s

Median Prefill TPS

898.33

Median Gen TPS

24.84

Model Information

Context Size

262144

Quantization

r64

Engine

vllm

Creation Method

LoRA Finetune

Model Type

Gemma31B

Chat Template

Gemma4

Reasoning

Yes

Vision

Yes

Parameters

31B

Added At

7/25/2026


language:

  • en license: gemma license_link: https://ai.google.dev/gemma/docs/gemma_4_license library_name: transformers pipeline_tag: text-generation base_model:
  • coder3101/gemma-4-31B-it-heretic base_model_relation: finetune tags:
  • roleplay
  • creative-writing
  • gemma
  • gemma4
  • qlora
  • uncensored
  • nsfw
  • not-for-all-audiences

Lilith

Lilith-31B-v1.0

Versatile uncensored roleplay / creative-writing model on Gemma-4-31B. Drives any character card (SillyTavern / pluma), built to be lively and varied rather than flat or repetitive. darthcrawl. Explicit-capable.

  • Base: coder3101/gemma-4-31B-it-heretic (decensored Gemma-4-31B-it)
  • Method: QLoRA r32/alpha64 all-linear, loss masked to model turns, 1 epoch eval-driven. Release = the ~0.39-epoch checkpoint (eval 2.11 vs base 7.03), picked over the fully-trained one to keep prose lively (loss != liveliness).
  • Data: ~30M tokens, 8 curated sources (human forum RP, AO3, curated public RP, curated synth), ~40/60 NSFW/SFW.
  • Chat template: Gemma-4 (<|turn>user ... <turn|> / <|turn>model ...).
  • Sampling tip: DRY + XTC + modest repetition penalty for max variety.

Siblings: Lilith-31B-v1.0 (bf16) | -LoRA | -GGUF | -MLX-4bit/6bit/8bit

Merged bf16 weights (~63 GB). For quantized inference use -GGUF or -MLX-*.