Skip to content
Hermit888Public

About

Compare SFT and DPO through blind human evaluation of RPG NPC quest dialogue.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

24 Commits

Folders and files

Repository files navigation

Overview

This project fine-tunes Llama-3.1-8B-Instruct to generate character-consistent NPC quest dialogues using a two-stage alignment pipeline (SFT → DPO). Given an NPC profile (name, background, location, quest), the model produces structured conversations covering an opening greeting, quest development, and a proactive farewell. A Gradio demo on HuggingFace Spaces enables blind side-by-side comparison of the SFT and DPO models.

Repository Structure

NPCAlign/
├── get_data/
│   ├── gather.py              # Download & split dataset (80/10/10)
│   ├── data_builder.py        # Parse system prompt → structured fields; normalise action format
│   ├── gen_DPO_data.py        # Generate SFT candidates, score with Gemma-4-26B, write DPO pairs
│   ├── split_DPO_dev.py       # Stratified train/dev split for DPO data
│   ├── orig_data/             # Raw JSONL splits (train / dev / test)
│   ├── pro_data/              # Processed JSONL splits used for SFT
│   ├── DPO_data/              # Raw DPO preference pairs (train.jsonl)
│   └── pro_DPO_data/          # Final DPO splits (pro_train / pro_dev)
│
├── models/
│   ├── sft_model.py           # SFT training (LoRA rank-16, max_length 3072)
│   ├── dpo_model.py           # DPO fine-tuning on SFT best checkpoint
│   ├── analyze_SFTtraining.py # Plot SFT training curves & metrics
│   ├── analyze_DPOtraining.py # Plot DPO training curves; compare with SFT baseline
│   ├── sft_results/           # SFT checkpoints, logs, test_results.json, plots
│   └── dpo_results/           # DPO checkpoints, logs, test_results.json, plots
│
├── demo/
│   ├── app.py                 # Gradio UI — 6-step blind evaluation flow
│   ├── inference.py           # ZeroGPU model loading + NPC response generation
│   ├── requirements.txt       # HuggingFace Space dependencies
│   └── README.md              # Space card (YAML frontmatter + usage notes)
│
└── README.md

Model

Method

Datasets

  • Training data: chimbiwide/NPC-Quest-Dialogue — 1,980 RPG quest conversations in total (13–19 turns each), split into 1,584 train / 198 dev / 198 test (80/10/10). Before training, all data were converted to a uniform format (*action* message).
  • Synthetic data: DPO preference pairs (chosen, rejected) generated by the SFT model and scored by Gemma 4 26B A4B. More details are in the Data Generation section. 1,341 pairs in total (opening: 454, development: 473, resolution: 414); 1,200 pairs for DPO training and 141 for DPO dev evaluation.

Implement for SFT

Input Format

SYSTEM_TEMPLATE = (
    "You are {name}.\n"
    "Background: {background}\n"
    "Current Location: {location}\n"
    "Quest: {quest}\n\n"
    "Roleplaying Instructions:\n"
    "- Speak using appropriate tone and vocabulary\n"
    "- Reference your background and current surroundings naturally\n"
    "- Keep responses conversational and authentic\n"
    "- React to the player's words and intentions\n"
    "- Your first response should be a greeting to the player\n"
    "- Once the quest has been fully explained and the player shows readiness, "
    "or if the conversation becomes repetitive, proactively bring it to a "
    "natural close with a farewell."
)

Training Configuration

Parameter Value
LoRA rank 16
LoRA alpha 32
LoRA target modules q/k/v/o projections + gate/up/down projections
Epochs 3
Learning rate 2e-4
Max length 3072 tokens
Loss masking Assistant turns only

Design Decisions

  • Loss masking is applied so that the training loss is calculated only on assistant turns, ignoring system and user turns.
  • LoRA is used instead of full-parameter fine-tuning to reduce computational cost.
  • max_length=3072 is set to cover 100% of conversations in the dataset without truncation, based on the token count of the longest conversation.
  • A farewell instruction is added to the system prompt to teach the model to proactively conclude conversations, extending the original dataset's roleplaying instructions.

Evaluation

  • During validation, only ROUGE-L and BERTScore-F1 are used for early stopping and best checkpoint selection, as they are faster to compute than BLEURT.
  • For the test set, evaluation includes ROUGE-L, BERTScore-F1, Self-BLEU, and BLEURT for a more comprehensive assessment. To analyse performance across conversation phases, one assistant turn is sampled from each of the three phases (opening, development, resolution) per test conversation. Final results are the mean across all three phases.

Convergence & Best Checkpoint

SFT Train Loss SFT Eval Loss SFT Eval Metrics The training loss decreases steadily. After step 400, the eval loss begins to increase slightly and ROUGE-L / BERTScore-F1 on the dev set plateau or decrease, indicating the onset of overfitting. The best checkpoint is therefore selected at step 400.

Data Generation

Purpose

The original dataset does not contain (chosen, rejected) pairs required for DPO training. Preference pairs are therefore synthesised using the SFT model and an LLM judge.

Process

  1. For each row in the train split, sample one assistant turn per phase (opening / development / resolution).
  2. For each sampled turn, build a ground-truth context prompt (system message + conversation history up to but not including the target turn) and generate 2 candidate responses using the SFT model at temperatures 0.7 and 1.0. Resolution-phase turns are sampled from the end of each conversation to ensure farewell behaviour is captured.
  3. Score all 3 candidates (original ground truth + 2 generated) in a single Gemma 4 26B API call using a 5-criterion rubric (character consistency, phase appropriateness, quest relevance, action format, and language quality; total 25 points).
  4. Set chosen = highest-scoring candidate, rejected = lowest-scoring. Skip pairs where score gap < MIN_SCORE_GAP (default 2 / 25).
  5. 1,341 pairs were retained after filtering (opening: 454, development: 473, resolution: 414).

Implement for DPO

Training Configuration

Parameter Value
Beta 0.1
Epochs 2
LoRA rank 16
LoRA alpha 32
Learning rate 5e-5
Max length 3072 tokens
Preference pairs 1,341

Design Decisions

  • The base model (Llama-3.1-8B-Instruct) is first merged with the best SFT LoRA checkpoint. This merged model serves as the frozen reference model, ensuring the KL constraint keeps the DPO model close to the already-learned conversation style.
  • Beta = 0.1 is the standard value that balances preference learning against the KL penalty from the reference model.

Evaluation

  • During validation, both reward margin and reward accuracy are computed on the dev set. Reward margin (avg chosen reward − avg rejected reward) is used as the criterion for best checkpoint selection, as it reflects the strength of the model's preference signal.
  • For test evaluation, the same four metrics used in SFT (ROUGE-L, BERTScore-F1, Self-BLEU, BLEURT) are reused to enable direct comparison.

Convergence & Best Checkpoint

DPO Train Loss DPO Reward Accuracy DPO Reward Margin DPO Reward Curves The training loss decreases and the reward margin on the dev set increases steadily, indicating convergence. The best checkpoint is selected at step 210 (out of 300), where the dev reward margin reaches its peak at 2.053.

Results

SFT

Phase ROUGE-L Self-BLEU BERTScore-F1 BLEURT
Opening 0.264 0.208 0.883 -0.692
Development 0.231 0.150 0.880 -0.719
Resolution 0.259 0.269 0.885 -0.719
Overall 0.251 0.264 0.883 -0.710

DPO

Phase ROUGE-L Self-BLEU BERTScore-F1 BLEURT
Opening 0.223 0.151 0.873 -0.843
Development 0.191 0.105 0.869 -0.811
Resolution 0.203 0.168 0.871 -0.867
Overall 0.206 0.187 0.871 -0.840

Plot of SFT vs. DPO

SFT vs. DPO

Analysis

  • ROUGE-L & Self-BLEU: Both decrease in the DPO model. Lower Self-BLEU indicates greater generation diversity; the corresponding ROUGE-L decrease is an expected trade-off, as more diverse outputs naturally diverge from the fixed reference answers.
  • BERTScore-F1: Both models score in the 0.87–0.88 range, indicating high semantic relevance to the reference answers is maintained after DPO.
  • BLEURT: The lower BLEURT for DPO (−0.840 vs −0.710) reflects that DPO responses deviate more from the reference answers in phrasing and content structure, consistent with the increased diversity observed in Self-BLEU.

Architecture of Demo

Important Process

  • Conversational turns: To enable the model to conclude conversations smoothly and proactively, two mechanisms are implemented. A "soft hint" system message is injected into the context from turn 10 onward, signalling the model to begin wrapping up; a hard maximum of 16 turns enforces an absolute limit on conversation length.
  • Evaluation: Users rate each model on a five-dimensional questionnaire (character consistency, dialogue fluency, quest completeness, ending quality, overall experience) on a scale of 1 to 5.

Demo

NPCAlign Demo NPC Quest Dialogue Demo

About

Compare SFT and DPO through blind human evaluation of RPG NPC quest dialogue.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages