Ziyu Luo · Wei Li · Haoping Huang · Kecheng Liu · Min Gao*
Chongqing University · Chongqing University Central Hospital · University of Reading
Official research code for the SharP article published in Expert Systems with Applications.
SharP is a boundary-aware augmentation framework for fake-news detection when labeled data are scarce. It combines a learnable soft prompt with a structured hard prompt, queries an LLM around uncertain classifier boundaries, and iteratively adds useful synthetic samples to the labeled set.
The repository provides dataset-specific workflows for PHEME, LIAR, and Twitter15/16.
Ziyu Luo, Wei Li, Haoping Huang, Kecheng Liu, and Min Gao. “SharP: Soft and Hard Prompt-Guided Augmentation with LLM for Low Resource Fake News Detection.” Expert Systems with Applications, Volume 309, Article 131178, 2026. https://doi.org/10.1016/j.eswa.2026.131178
- Available online: 10 January 2026
- Volume publication date: 5 May 2026
- Volume: 309
- Article number: 131178
- PII: S0957417426000928
The repository copy at paper/SharP.pdf is the author manuscript retained with the research artifact. Please use the DOI above to access and cite the publisher's Version of Record.
- Dual-prompt generation: hard prompts supply task instructions while soft prompts learn dataset-specific semantic patterns.
- Boundary-aware querying: entropy-based active selection focuses augmentation on samples near the detector's decision boundary.
- Detector-guided prompt optimization: classifier feedback and representation alignment guide soft-prompt learning.
- Iterative low-resource learning: selected and generated samples extend the labeled set over multiple active-learning rounds.
- Three benchmark families: preprocessing and experiment workflows for PHEME, LIAR, and Twitter15/16.
Figure 1. At round t, SharP scores the unlabeled pool, queries boundary samples, augments them with the prompt-guided LLM, and extends the labeled set for round t+1.
Given a labeled set
train detector on D_L^t
↓
score D_U^t and query high-entropy samples S_t
↓
optimize soft prompt with detector feedback
↓
combine soft prompt + hard prompt for LLM augmentation
↓
D_L^(t+1) = D_L^t ∪ S'_t
D_U^(t+1) = D_U^t \ S_t
Figure 2. SharP first pre-trains the detector and prompt module, then alternates boundary querying, prompt-guided augmentation, and detector fine-tuning.
The detector combines:
- a BERT text encoder;
- TextCNN branches for local semantic patterns;
- BiGRU branches over handcrafted affective features;
- a fake/real classifier;
- event-domain and marked/unmarked auxiliary heads.
The augmentation module freezes the causal LLM backbone and optimizes a 25-token soft prompt. During fine-tuning, selected pool samples are augmented and used to refine the detector boundary.
Figure 3. The hard prompt Th and soft prompt Ts are combined with the input embedding. Detector feedback aligns the generated representation with task-relevant features.
The augmentation path is implemented primarily in DAAL/amodel.py, while detector training, active selection, and iterative data extension are implemented in DAAL/model_adjust.py.
SharP/
├── DAAL/ # Core SharP model and training loop
│ ├── amodel.py # Soft/hard prompt LLM module
│ ├── model_adjust.py # Detector + active augmentation loop
│ ├── process_data.py # PHEME-oriented tensor conversion
│ └── run_sweep.py # Low-resource label-budget sweep
├── idea1_features/ # Affective/linguistic feature code
├── experiments/
│ ├── pheme/ # Retained PHEME experiment version
│ └── cross_benchmark/ # LIAR/Twitter preprocessing and adaptation
├── data/README.md # Upstream sources and expected local layout
├── assets/figures/ # Paper architecture figures
├── scripts/smoke_check.py # Offline structure/data audit
├── docs/ # Reproduction and version notes
└── paper/SharP.pdf # Author manuscript
Follow data/README.md to obtain the upstream datasets and lexicons, prepare tensors and prompt seeds, and create the expected local layout.
conda create -n sharp python=3.10 -y
conda activate sharp
pip install -r requirements.txtSet a BERT-compatible encoder and a causal LLM:
export SHARP_BERT_MODEL=bert-base-uncased
export SHARP_LLM_MODEL=/path/to/zephyr-7b-beta-awqPowerShell:
$env:SHARP_BERT_MODEL = "bert-base-uncased"
$env:SHARP_LLM_MODEL = "D:\models\zephyr-7b-beta-awq"For an AWQ checkpoint, install the loader required by that checkpoint. A CUDA-capable GPU is strongly recommended for the full BERT + 7B-LLM pipeline.
From the repository root:
python scripts/smoke_check.pyWithout external datasets, the check validates Python syntax and the main experiment entry points. After the expected data files are prepared, it additionally verifies dataset totals, the PHEME event mapping, all five PHEME pools, and exact clippool membership.
event_label |
Event | Raw file |
|---|---|---|
| 0 | Charlie Hebdo | data/pheme/raw/charliehebdo.csv |
| 1 | Ferguson | data/pheme/raw/ferguson.csv |
| 2 | Germanwings crash | data/pheme/raw/germanwings-crash.csv |
| 3 | Ottawa shooting | data/pheme/raw/ottawashooting.csv |
| 4 | Sydney siege | data/pheme/raw/sydneysiege.csv |
The mapping is confirmed by both the ordered input list in add_event_label(...) and the values stored in the raw CSV files.
data/pheme/prompt_seed/clippool_16.xlsx contains 19 rows from the Germanwings event-2 pool. The loader uses batch_size=16, drop_last=True, producing one 16-sample pre-training batch.
For LIAR and Twitter15/16, python scripts/build_release_prompt_seeds.py deterministically samples 19-row seeds from each dataset's local pool with random_state=42.
Historical code uses working-directory-relative paths. Run each command from the directory shown below.
This is the retained configuration with mutually consistent train, pool, test, and prompt-seed artifacts.
export SHARP_PHEME_EVENT=2
export SHARP_PROMPT_SEED=../../data/pheme/prompt_seed/clippool_16.xlsx
cd experiments/pheme
python model_adjust.pyPowerShell:
cd experiments\pheme
$env:SHARP_PHEME_EVENT = "2"
$env:SHARP_PROMPT_SEED = "..\..\data\pheme\prompt_seed\clippool_16.xlsx"
python model_adjust.pyThe preparation workflow supports per-event pool/test/validation artifacts for event ids 0–4. Rebuild the shared source/train tensors and prompt seed for the selected target event.
The main runner uses the retained LIAR tensors and supports the paper's low-resource ratio argument:
cd DAAL
python model_adjust.py \
--less-frac 0.10 \
--llm-model "$SHARP_LLM_MODEL"--less-frac is the labeled-data ratio. The paper evaluates:
1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 70%
Run the complete sweep after setting a LIAR-specific seed:
cd DAAL
python run_sweep.pyThe default is data/liar/prompt_seed/clippool_16_release.xlsx. Use --prompt-seed or SHARP_PROMPT_SEED to select another seed file.
After obtaining the datasets from their upstream sources, rebuild clustered event/domain labels and tensors locally:
python experiments/cross_benchmark/twitter_processing.py
python experiments/cross_benchmark/process_data.pyRun the retained Twitter adaptation:
cd experiments/cross_benchmark
python model_adjust.py \
--less-frac 0.05 \
--llm-model "$SHARP_LLM_MODEL"The runner enables both prompt-guided generated-text insertion points and defaults to data/twitter15_16/prompt_seed/clippool_16_release.xlsx.
For a single low-resource run, SharP performs the following stages:
- Load data: BERT token ids, masks, handcrafted 8×24 affective features, class labels, event labels, and active-pool indices.
- Build the detector: initialize BERT, TextCNN, BiGRU, the fake/real head, and auxiliary domain heads.
- Low-resource sampling: retain the requested fraction of labeled training data while keeping validation and test sets fixed.
- Detector pre-training: optimize the fake-news detector and auxiliary objectives.
- Boundary query: rank unlabeled samples by predictive entropy and select uncertain samples.
- Soft-prompt pre-training: initialize and train the 25-token prompt using the dataset-specific seed set.
- Prompt-guided generation: combine the learned soft prefix with the hard rewrite prompt and generate augmented text.
- Set extension: add selected/generated samples to the labeled set and remove queried samples from the pool.
- Fine-tuning: update the detector and soft prompt, evaluate, and repeat until the pool is exhausted or querying stops.
Runtime logs and checkpoints are written to ignored output directories.
| Component | Location |
|---|---|
| Core detector and prompt module | DAAL/ |
| PHEME preparation and experiment workflow | experiments/pheme/ |
| LIAR and Twitter15/16 workflows | experiments/cross_benchmark/ |
| Prompt-seed generation | scripts/build_release_prompt_seeds.py |
| Data preparation and checks | data/README.md, scripts/smoke_check.py |
For dataset and workflow details, see docs/VERSION_SELECTION.md, docs/DATA_SUMMARY.md, and the 中文运行清单.
If this repository is useful in your research, please cite:
@article{luo2026sharp,
title = {SharP: Soft and Hard Prompt-Guided Augmentation with LLM for Low Resource Fake News Detection},
author = {Luo, Ziyu and Li, Wei and Huang, Haoping and Liu, Kecheng and Gao, Min},
journal = {Expert Systems with Applications},
volume = {309},
pages = {131178},
year = {2026},
doi = {10.1016/j.eswa.2026.131178},
url = {https://doi.org/10.1016/j.eswa.2026.131178}
}Machine-readable citation metadata is available in CITATION.cff.
Dataset and lexicon terms are independent of the source code. See docs/DATA_POLICY.md.


