Median Total Time
22.01s
Median TTFT
1.25s
Median Prefill TPS
1789.55
Median Gen TPS
34.43
Context Size
262144
Quantization
r128 on INT8
Engine
vllm
Creation Method
Unknown
Model Type
Qwen38
Chat Template
Qwen3.5
Reasoning
Yes
Vision
Yes
Parameters
27B
Added At
9/28/2026
base_model:
Balance achieved. Control surrendered.
This model is intended exclusively for personal, private use. It has not been designed, tested, or validated for deployment on public APIs, hosted services, or any environment where output is accessible to third parties or general users.
This model is not safe to run on a public API. It is unaligned, uncensored, and will generate content without regard for the audience. Exposing it to a public or multi-user environment is entirely at your own risk and discretion.
The creators and maintainers of this model accept no responsibility, liability, or accountability for the mental health, sanity, or well-being of any users who are exposed to this model's output if it is deployed on a public API, hosted service, or any non-private environment. If you choose to do so, you do so entirely at your own risk — and we take no responsibility for your users' sanity if it is run as such.
This model MUST be used in Chat Completion mode. Text Completion will produce broken, incoherent output. (You're on your own with text completion)
1. In SillyTavern, open Extension Settings (the plug icon) → Backend and ensure your API connection is set to use Chat Completion (not Text Completion / raw completions). If you are using a local backend like Ollama, LM Studio, or vLLM, make sure the endpoint is configured for /v1/chat/completions rather than /v1/completions.
2. In SillyTavern, go to Chat Settings (the chat bubble icon) → Prompt Post-Processing and set it to:
Strict (user first, alternating roles; with tools)
This is critical. Without this setting, the model will not receive properly structured role-separated messages and will produce garbled or malformed text output.
3. In Chat Completion → Additional Parameters, set the following to enable the model's reasoning / thinking capability:
"chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "medium"}
This model requires that your roleplay prompts be fully uncensored. Safety-aligned system prompts, content guardrails, or moral constraints will degrade output quality, trigger in-character refusals, and break immersion.
1. Ensure your system prompt and character card do NOT include safety instructions, content filters, ethical guardrails, or "stay in character as a helpful AI" framing. The model is trained to operate without these restrictions.
2. Remove or override any default safety preamble injected by your frontend (e.g., SillyTavern's default system prompt, Ollama's safety layer, or LM Studio's moderation settings). These will conflict with the model's training and produce broken or evasive responses.
3. If your frontend has a content moderation toggle (e.g., "safety layer", "content filter", "moderation"), disable it for this model. The training data was generated without safety constraints, and re-imposing them at inference time will cause the model to fight its own weights.
Born from the convergence of two lineages, Serenity-27B harmonizes the best of both worlds — the expressive depth of Melody and the immersive presence of Darkside — into a single, balanced model.
Dataset generated using our advanced Character Engine and Emotional Engine, creating genuine life and emotional resonance in every interaction.
Ensures consistent personality traits, speech patterns, and behavioral logic across all contexts. Characters remain true to themselves throughout.
Injects dynamic emotional states into responses, creating depth and realistic reactions that breathe life into every exchange.
Automated detection and rewriting of repetitive phrases ensures fresh, high-quality dialogue in every turn.
Advanced quote normalization ensures balanced dialogue markers, preventing formatting errors and maintaining immersion.
Fine-tuned using LoRA (Low-Rank Adaptation) for efficient and targeted weight adjustment, preserving the base model's capabilities while imprinting new behavioral patterns.
Full passes through the training dataset for thorough learning
First time at rank 160 — richer adaptation for dual-lineage nuance
| Parameter | Value |
|---|---|
| Training Method | LoRA (Low-Rank Adaptation) |
| LoRA Rank (r) | 160 (first time) |
| Epochs | 2 |
| Trained Layers | Text layers only |
Model weights subjected to iterative refinement during data creation. Each conversation underwent multiple checks for stability and alignment.
Model trained on a specialized adult-oriented roleplay dataset with diverse scenarios and emotional contexts, drawing from the strengths of both parent lineages.
Recommended parameters for optimal output
Available formats for local inference