How much of the Fediverse was not written by a person, and how far does it travel?
University project (Social Networks and Media), developed with Paolo Pangallo and Constantin Adrian Antoci, that collects posts from Mastodon, estimates which ones were written by an AI, checks verifiable claims and studies how content spreads through the social network. The results are explored through a web interface. Code, comments and documentation are in Italian.
- Collection. For each topic in
topic_list.txtit finds the most active Mastodon instances, discovers their most used hashtags and downloads the posts, together with accounts and network edges (follows, boosts, replies). Everything goes into a database (PostgreSQL or SQLite). - Synthetic-text detection. Every text goes through independent detectors: Fast-DetectGPT, AdaDetectGPT, Binoculars and Desklib. A comparison across detectors (
comparatore_detector/) shows where they agree and where they do not. - Verification. A model estimates whether a post contains a checkable claim (check-worthiness); claims are then verified by an LLM that searches the web for evidence and cites its sources.
- Network and diffusion.
- Community detection on the social graph (Leiden, Infomap).
- Influence maximization with CELF++, PMIA and SKIM, to choose the nodes to start from.
- Monte Carlo simulations (Independent Cascade model) and TwitterRank to measure how far a piece of content travels.
- Interface. A FastAPI backend and a React frontend that present the work as a sequence of chapters: the corpora, synthetic text, verification, propagation.
| Path | Content |
|---|---|
pipeline.py, snm/collection/ |
Data collection from Mastodon |
snm/storage/, db/ |
Database access and schemas (PostgreSQL, SQLite) |
snm/analysis/ |
Text export, check-worthiness, fact-checking, DB import/export |
snm/graph/ |
Graph construction, community detection, visualizations |
binoculars/, desklib_detector/, comparatore_detector/ |
Synthetic-text detectors and comparison |
Max_Influence/ |
Influence-maximization algorithms and results |
misinformation_impact/ |
Misinformation diffusion simulations |
webapp/ |
FastAPI API |
frontend/ |
React + Vite interface |
tests/ |
Backend tests (pytest) |
For a per-file description see report.md; for the product idea, PRODUCT.md and DESIGN.md (in Italian).
- Python 3.10+
- Node.js 18+
- A database: SQLite (no installation) or PostgreSQL
- Optional: a GPU for the model-based detectors (
torch,transformers)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cd frontend && npm install && cd ..
cp env.example .env # then fill in the valuesIn .env set at least DATABASE_URL. INSTANCES_SOCIAL_API is needed to discover instances, and the MASTODON_TOKEN_* tokens are optional: without them anonymous access is used, with a lower rate limit.
Everything at once on Windows:
.\start_all.ps1Or separately:
uvicorn webapp.main:app --port 8088 --reload # backend → http://127.0.0.1:8088
cd frontend && npm run dev # frontend → http://localhost:5173To collect new data:
python pipeline.pyThe repository does not contain the post corpus and the heavy results (post_texts.jsonl, the full detector scores, the fact-checking reports): they are content published by other people, and they can be regenerated with the pipeline. Small aggregated results stay in the repository (Max_Influence/Risultati_IM/, Risultati_Binoculars/, the ROC plots in data/).
For how output files are looked up and how to customize their paths see DATASET_SETUP.md.
pip install pytest
pytest tests/
cd frontend && npm test