This project fine-tunes Llama-3.1-8B-Instruct to generate character-consistent NPC quest dialogues using a two-stage alignment pipeline (SFT → DPO). Given an NPC profile (name, background, location, quest), the model produces structured conversations covering an opening greeting, quest development, and a proactive farewell. A Gradio demo on HuggingFace Spaces enables blind side-by-side comparison of the SFT and DPO models.
NPCAlign/
├── get_data/
│ ├── gather.py # Download & split dataset (80/10/10)
│ ├── data_builder.py # Parse system prompt → structured fields; normalise action format
│ ├── gen_DPO_data.py # Generate SFT candidates, score with Gemma-4-26B, write DPO pairs
│ ├── split_DPO_dev.py # Stratified train/dev split for DPO data
│ ├── orig_data/ # Raw JSONL splits (train / dev / test)
│ ├── pro_data/ # Processed JSONL splits used for SFT
│ ├── DPO_data/ # Raw DPO preference pairs (train.jsonl)
│ └── pro_DPO_data/ # Final DPO splits (pro_train / pro_dev)
│
├── models/
│ ├── sft_model.py # SFT training (LoRA rank-16, max_length 3072)
│ ├── dpo_model.py # DPO fine-tuning on SFT best checkpoint
│ ├── analyze_SFTtraining.py # Plot SFT training curves & metrics
│ ├── analyze_DPOtraining.py # Plot DPO training curves; compare with SFT baseline
│ ├── sft_results/ # SFT checkpoints, logs, test_results.json, plots
│ └── dpo_results/ # DPO checkpoints, logs, test_results.json, plots
│
├── demo/
│ ├── app.py # Gradio UI — 6-step blind evaluation flow
│ ├── inference.py # ZeroGPU model loading + NPC response generation
│ ├── requirements.txt # HuggingFace Space dependencies
│ └── README.md # Space card (YAML frontmatter + usage notes)
│
└── README.md
- Base model: meta-llama/Meta-Llama-3.1-8B-Instruct
- SFT adapter: HermitQ/NPCAlign-SFT
- DPO adapter: HermitQ/NPCAlign-DPO
- Training data: chimbiwide/NPC-Quest-Dialogue — 1,980 RPG quest conversations in total (13–19 turns each), split into 1,584 train / 198 dev / 198 test (80/10/10). Before training, all data were converted to a uniform format (*action* message).
- Synthetic data: DPO preference pairs (chosen, rejected) generated by the SFT model and scored by Gemma 4 26B A4B. More details are in the Data Generation section. 1,341 pairs in total (opening: 454, development: 473, resolution: 414); 1,200 pairs for DPO training and 141 for DPO dev evaluation.
SYSTEM_TEMPLATE = (
"You are {name}.\n"
"Background: {background}\n"
"Current Location: {location}\n"
"Quest: {quest}\n\n"
"Roleplaying Instructions:\n"
"- Speak using appropriate tone and vocabulary\n"
"- Reference your background and current surroundings naturally\n"
"- Keep responses conversational and authentic\n"
"- React to the player's words and intentions\n"
"- Your first response should be a greeting to the player\n"
"- Once the quest has been fully explained and the player shows readiness, "
"or if the conversation becomes repetitive, proactively bring it to a "
"natural close with a farewell."
)
| Parameter | Value |
|---|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA target modules | q/k/v/o projections + gate/up/down projections |
| Epochs | 3 |
| Learning rate | 2e-4 |
| Max length | 3072 tokens |
| Loss masking | Assistant turns only |
- Loss masking is applied so that the training loss is calculated only on assistant turns, ignoring system and user turns.
- LoRA is used instead of full-parameter fine-tuning to reduce computational cost.
max_length=3072is set to cover 100% of conversations in the dataset without truncation, based on the token count of the longest conversation.- A farewell instruction is added to the system prompt to teach the model to proactively conclude conversations, extending the original dataset's roleplaying instructions.
- During validation, only ROUGE-L and BERTScore-F1 are used for early stopping and best checkpoint selection, as they are faster to compute than BLEURT.
- For the test set, evaluation includes ROUGE-L, BERTScore-F1, Self-BLEU, and BLEURT for a more comprehensive assessment. To analyse performance across conversation phases, one assistant turn is sampled from each of the three phases (opening, development, resolution) per test conversation. Final results are the mean across all three phases.
The training loss decreases steadily. After step 400, the eval loss begins to increase slightly and ROUGE-L / BERTScore-F1 on the dev set plateau or decrease, indicating the onset of overfitting. The best checkpoint is therefore selected at step 400.
The original dataset does not contain (chosen, rejected) pairs required for DPO training. Preference pairs are therefore synthesised using the SFT model and an LLM judge.
- For each row in the train split, sample one assistant turn per phase (opening / development / resolution).
- For each sampled turn, build a ground-truth context prompt (system message + conversation history up to but not including the target turn) and generate 2 candidate responses using the SFT model at temperatures 0.7 and 1.0. Resolution-phase turns are sampled from the end of each conversation to ensure farewell behaviour is captured.
- Score all 3 candidates (original ground truth + 2 generated) in a single Gemma 4 26B API call using a 5-criterion rubric (character consistency, phase appropriateness, quest relevance, action format, and language quality; total 25 points).
- Set chosen = highest-scoring candidate, rejected = lowest-scoring. Skip pairs where score gap < MIN_SCORE_GAP (default 2 / 25).
- 1,341 pairs were retained after filtering (opening: 454, development: 473, resolution: 414).
| Parameter | Value |
|---|---|
| Beta | 0.1 |
| Epochs | 2 |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| Learning rate | 5e-5 |
| Max length | 3072 tokens |
| Preference pairs | 1,341 |
- The base model (Llama-3.1-8B-Instruct) is first merged with the best SFT LoRA checkpoint. This merged model serves as the frozen reference model, ensuring the KL constraint keeps the DPO model close to the already-learned conversation style.
- Beta = 0.1 is the standard value that balances preference learning against the KL penalty from the reference model.
- During validation, both reward margin and reward accuracy are computed on the dev set. Reward margin (avg chosen reward − avg rejected reward) is used as the criterion for best checkpoint selection, as it reflects the strength of the model's preference signal.
- For test evaluation, the same four metrics used in SFT (ROUGE-L, BERTScore-F1, Self-BLEU, BLEURT) are reused to enable direct comparison.
The training loss decreases and the reward margin on the dev set increases steadily, indicating convergence. The best checkpoint is selected at step 210 (out of 300), where the dev reward margin reaches its peak at 2.053.
| Phase | ROUGE-L | Self-BLEU | BERTScore-F1 | BLEURT |
|---|---|---|---|---|
| Opening | 0.264 | 0.208 | 0.883 | -0.692 |
| Development | 0.231 | 0.150 | 0.880 | -0.719 |
| Resolution | 0.259 | 0.269 | 0.885 | -0.719 |
| Overall | 0.251 | 0.264 | 0.883 | -0.710 |
| Phase | ROUGE-L | Self-BLEU | BERTScore-F1 | BLEURT |
|---|---|---|---|---|
| Opening | 0.223 | 0.151 | 0.873 | -0.843 |
| Development | 0.191 | 0.105 | 0.869 | -0.811 |
| Resolution | 0.203 | 0.168 | 0.871 | -0.867 |
| Overall | 0.206 | 0.187 | 0.871 | -0.840 |
- ROUGE-L & Self-BLEU: Both decrease in the DPO model. Lower Self-BLEU indicates greater generation diversity; the corresponding ROUGE-L decrease is an expected trade-off, as more diverse outputs naturally diverge from the fixed reference answers.
- BERTScore-F1: Both models score in the 0.87–0.88 range, indicating high semantic relevance to the reference answers is maintained after DPO.
- BLEURT: The lower BLEURT for DPO (−0.840 vs −0.710) reflects that DPO responses deviate more from the reference answers in phrasing and content structure, consistent with the increased diversity observed in Self-BLEU.
- Conversational turns: To enable the model to conclude conversations smoothly and proactively, two mechanisms are implemented. A "soft hint" system message is injected into the context from turn 10 onward, signalling the model to begin wrapping up; a hard maximum of 16 turns enforces an absolute limit on conversation length.
- Evaluation: Users rate each model on a five-dimensional questionnaire (character consistency, dialogue fluency, quest completeness, ending quality, overall experience) on a scale of 1 to 5.
NPCAlign Demo NPC Quest Dialogue Demo
