Skip to content

Repository files navigation

SharP

Soft and Hard Prompt-Guided Augmentation with LLM for Low-Resource Fake News Detection

Ziyu Luo · Wei Li · Haoping Huang · Kecheng Liu · Min Gao*

Chongqing University · Chongqing University Central Hospital · University of Reading

Python PyTorch Transformers Code Check DOI Release

Official research code for the SharP article published in Expert Systems with Applications.

TL;DR

SharP is a boundary-aware augmentation framework for fake-news detection when labeled data are scarce. It combines a learnable soft prompt with a structured hard prompt, queries an LLM around uncertain classifier boundaries, and iteratively adds useful synthetic samples to the labeled set.

The repository provides dataset-specific workflows for PHEME, LIAR, and Twitter15/16.

Publication

Ziyu Luo, Wei Li, Haoping Huang, Kecheng Liu, and Min Gao. “SharP: Soft and Hard Prompt-Guided Augmentation with LLM for Low Resource Fake News Detection.” Expert Systems with Applications, Volume 309, Article 131178, 2026. https://doi.org/10.1016/j.eswa.2026.131178

  • Available online: 10 January 2026
  • Volume publication date: 5 May 2026
  • Volume: 309
  • Article number: 131178
  • PII: S0957417426000928

The repository copy at paper/SharP.pdf is the author manuscript retained with the research artifact. Please use the DOI above to access and cite the publisher's Version of Record.

Highlights

  • Dual-prompt generation: hard prompts supply task instructions while soft prompts learn dataset-specific semantic patterns.
  • Boundary-aware querying: entropy-based active selection focuses augmentation on samples near the detector's decision boundary.
  • Detector-guided prompt optimization: classifier feedback and representation alignment guide soft-prompt learning.
  • Iterative low-resource learning: selected and generated samples extend the labeled set over multiple active-learning rounds.
  • Three benchmark families: preprocessing and experiment workflows for PHEME, LIAR, and Twitter15/16.

Framework Overview

SharP iterative active-learning framework

Figure 1. At round t, SharP scores the unlabeled pool, queries boundary samples, augments them with the prompt-guided LLM, and extends the labeled set for round t+1.

Given a labeled set $D_L^t$ and an unlabeled pool $D_U^t$, SharP alternates between detector training, uncertainty scoring, sample querying, LLM augmentation, and labeled-set extension:

train detector on D_L^t
        ↓
score D_U^t and query high-entropy samples S_t
        ↓
optimize soft prompt with detector feedback
        ↓
combine soft prompt + hard prompt for LLM augmentation
        ↓
D_L^(t+1) = D_L^t ∪ S'_t
D_U^(t+1) = D_U^t \ S_t

Model Architecture

SharP pre-training and fine-tuning architecture

Figure 2. SharP first pre-trains the detector and prompt module, then alternates boundary querying, prompt-guided augmentation, and detector fine-tuning.

The detector combines:

  • a BERT text encoder;
  • TextCNN branches for local semantic patterns;
  • BiGRU branches over handcrafted affective features;
  • a fake/real classifier;
  • event-domain and marked/unmarked auxiliary heads.

The augmentation module freezes the causal LLM backbone and optimizes a 25-token soft prompt. During fine-tuning, selected pool samples are augmented and used to refine the detector boundary.

Prompt-Guided Augmentation

Soft and hard prompt-guided LLM augmentation

Figure 3. The hard prompt Th and soft prompt Ts are combined with the input embedding. Detector feedback aligns the generated representation with task-relevant features.

The augmentation path is implemented primarily in DAAL/amodel.py, while detector training, active selection, and iterative data extension are implemented in DAAL/model_adjust.py.

Repository Layout

SharP/
├── DAAL/                         # Core SharP model and training loop
│   ├── amodel.py                 # Soft/hard prompt LLM module
│   ├── model_adjust.py           # Detector + active augmentation loop
│   ├── process_data.py           # PHEME-oriented tensor conversion
│   └── run_sweep.py              # Low-resource label-budget sweep
├── idea1_features/               # Affective/linguistic feature code
├── experiments/
│   ├── pheme/                    # Retained PHEME experiment version
│   └── cross_benchmark/          # LIAR/Twitter preprocessing and adaptation
├── data/README.md                # Upstream sources and expected local layout
├── assets/figures/               # Paper architecture figures
├── scripts/smoke_check.py        # Offline structure/data audit
├── docs/                         # Reproduction and version notes
└── paper/SharP.pdf               # Author manuscript

Follow data/README.md to obtain the upstream datasets and lexicons, prepare tensors and prompt seeds, and create the expected local layout.

Installation

1. Create the environment

conda create -n sharp python=3.10 -y
conda activate sharp
pip install -r requirements.txt

2. Configure pretrained models

Set a BERT-compatible encoder and a causal LLM:

export SHARP_BERT_MODEL=bert-base-uncased
export SHARP_LLM_MODEL=/path/to/zephyr-7b-beta-awq

PowerShell:

$env:SHARP_BERT_MODEL = "bert-base-uncased"
$env:SHARP_LLM_MODEL = "D:\models\zephyr-7b-beta-awq"

For an AWQ checkpoint, install the loader required by that checkpoint. A CUDA-capable GPU is strongly recommended for the full BERT + 7B-LLM pipeline.

3. Validate the release before training

From the repository root:

python scripts/smoke_check.py

Without external datasets, the check validates Python syntax and the main experiment entry points. After the expected data files are prepared, it additionally verifies dataset totals, the PHEME event mapping, all five PHEME pools, and exact clippool membership.

Data Preparation

PHEME event mapping

event_label Event Raw file
0 Charlie Hebdo data/pheme/raw/charliehebdo.csv
1 Ferguson data/pheme/raw/ferguson.csv
2 Germanwings crash data/pheme/raw/germanwings-crash.csv
3 Ottawa shooting data/pheme/raw/ottawashooting.csv
4 Sydney siege data/pheme/raw/sydneysiege.csv

The mapping is confirmed by both the ordered input list in add_event_label(...) and the values stored in the raw CSV files.

Soft-prompt seed

data/pheme/prompt_seed/clippool_16.xlsx contains 19 rows from the Germanwings event-2 pool. The loader uses batch_size=16, drop_last=True, producing one 16-sample pre-training batch.

For LIAR and Twitter15/16, python scripts/build_release_prompt_seeds.py deterministically samples 19-row seeds from each dataset's local pool with random_state=42.

Running SharP

Historical code uses working-directory-relative paths. Run each command from the directory shown below.

A. PHEME — retained event-2/Germanwings snapshot

This is the retained configuration with mutually consistent train, pool, test, and prompt-seed artifacts.

export SHARP_PHEME_EVENT=2
export SHARP_PROMPT_SEED=../../data/pheme/prompt_seed/clippool_16.xlsx
cd experiments/pheme
python model_adjust.py

PowerShell:

cd experiments\pheme
$env:SHARP_PHEME_EVENT = "2"
$env:SHARP_PROMPT_SEED = "..\..\data\pheme\prompt_seed\clippool_16.xlsx"
python model_adjust.py

The preparation workflow supports per-event pool/test/validation artifacts for event ids 0–4. Rebuild the shared source/train tensors and prompt seed for the selected target event.

B. LIAR — retained complete augmentation runner

The main runner uses the retained LIAR tensors and supports the paper's low-resource ratio argument:

cd DAAL
python model_adjust.py \
  --less-frac 0.10 \
  --llm-model "$SHARP_LLM_MODEL"

--less-frac is the labeled-data ratio. The paper evaluates:

1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 70%

Run the complete sweep after setting a LIAR-specific seed:

cd DAAL
python run_sweep.py

The default is data/liar/prompt_seed/clippool_16_release.xlsx. Use --prompt-seed or SHARP_PROMPT_SEED to select another seed file.

C. Twitter15/16

After obtaining the datasets from their upstream sources, rebuild clustered event/domain labels and tensors locally:

python experiments/cross_benchmark/twitter_processing.py
python experiments/cross_benchmark/process_data.py

Run the retained Twitter adaptation:

cd experiments/cross_benchmark
python model_adjust.py \
  --less-frac 0.05 \
  --llm-model "$SHARP_LLM_MODEL"

The runner enables both prompt-guided generated-text insertion points and defaults to data/twitter15_16/prompt_seed/clippool_16_release.xlsx.

Training Procedure

For a single low-resource run, SharP performs the following stages:

  1. Load data: BERT token ids, masks, handcrafted 8×24 affective features, class labels, event labels, and active-pool indices.
  2. Build the detector: initialize BERT, TextCNN, BiGRU, the fake/real head, and auxiliary domain heads.
  3. Low-resource sampling: retain the requested fraction of labeled training data while keeping validation and test sets fixed.
  4. Detector pre-training: optimize the fake-news detector and auxiliary objectives.
  5. Boundary query: rank unlabeled samples by predictive entropy and select uncertain samples.
  6. Soft-prompt pre-training: initialize and train the 25-token prompt using the dataset-specific seed set.
  7. Prompt-guided generation: combine the learned soft prefix with the hard rewrite prompt and generate augmented text.
  8. Set extension: add selected/generated samples to the labeled set and remove queried samples from the pool.
  9. Fine-tuning: update the detector and soft prompt, evaluate, and repeat until the pool is exhausted or querying stops.

Runtime logs and checkpoints are written to ignored output directories.

Repository Components

Component Location
Core detector and prompt module DAAL/
PHEME preparation and experiment workflow experiments/pheme/
LIAR and Twitter15/16 workflows experiments/cross_benchmark/
Prompt-seed generation scripts/build_release_prompt_seeds.py
Data preparation and checks data/README.md, scripts/smoke_check.py

For dataset and workflow details, see docs/VERSION_SELECTION.md, docs/DATA_SUMMARY.md, and the 中文运行清单.

Citation

If this repository is useful in your research, please cite:

@article{luo2026sharp,
  title   = {SharP: Soft and Hard Prompt-Guided Augmentation with LLM for Low Resource Fake News Detection},
  author  = {Luo, Ziyu and Li, Wei and Huang, Haoping and Liu, Kecheng and Gao, Min},
  journal = {Expert Systems with Applications},
  volume  = {309},
  pages   = {131178},
  year    = {2026},
  doi     = {10.1016/j.eswa.2026.131178},
  url     = {https://doi.org/10.1016/j.eswa.2026.131178}
}

Machine-readable citation metadata is available in CITATION.cff.

License and Data Policy

Dataset and lexicon terms are independent of the source code. See docs/DATA_POLICY.md.


SharP research-code release · public artifact

About

Official code for SharP, published in Expert Systems with Applications (Volume 309, Article 131178).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages