Skip to content

Latest commit

 

History

148 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Machine Learning Classification of Neutron Star Composition

Public code associated with Ioannis Papathanasiou's undergraduate thesis-era work on a restricted synthetic comparison of hadronic-star and self-bound strange-quark-star models.

Historical scope. This repository is a public code record associated with the project; it is not the examined thesis document or a verified archival copy of the exact code package examined in February 2026. A later redevelopment identified important limits on how the classifier outputs can be interpreted. See the portfolio case study and the classification risk audit.

Scientific interpretation boundary

The workflow studies model discrimination inside a constructed synthetic design. It does not measure the composition of real compact stars, establish general hadronic-versus-quark classification, or produce astrophysical posterior probabilities. Scores produced for named objects by the historical analyze_candidates.py script are model-dependent outputs from that synthetic design. They should not be read as observational findings or composition probabilities.

The later audit found that the effective number and diversity of independent physical families, rather than the number of generated stellar rows, controls the strength of the validation. Reported row-level accuracy, AUC, calibration, and named-object scores therefore do not by themselves establish generalization to unseen equation-of-state families or observations.

Project question and workflow

The project explored whether mass, radius, tidal deformability, and selected curve-derived features could separate two synthetic classes whose mass-radius sequences may overlap. The code:

  1. constructs a configurable library of synthetic hadronic and CFL quark-matter equations of state;
  2. solves stellar-structure and tidal equations to generate mass-radius-deformability sequences;
  3. trains Random Forest classifiers on several feature sets;
  4. separates equation-of-state curves between train and test sets using Curve_ID; and
  5. produces diagnostic and scientific visualizations.

The current source defines 20 active analytic hadronic core fits; a commented-out PS fit is not active. The default orchestration requests 20,000 generated curves and balances the resulting stellar rows before training. Those counts describe configured synthetic generation, not 20,000 independent physical theories or observational samples.

Repository structure

Bachelor_Thesis/
├── main.py                   # Generation, training, diagnostics, and plotting orchestration
├── data/                     # Generated datasets when created locally
├── plots/                    # Versioned and locally generated figures
└── src/
    ├── const.py              # Numerical settings, parameter ranges, and data schema
    ├── physics/              # EoS construction and stellar-structure solvers
    ├── ml_pipeline/          # Random Forest training, validation, and historical candidate script
    └── visualize/            # Static diagnostic and scientific plots

The repository also retains notebooks associated with several Python modules. The Python files are the clearer entry points for inspection.

Local use

Python 3.8 or later is the historical target. There is no lockfile, automated test suite, or reproducible environment specification in this repository, so a successful run is not guaranteed across current dependency versions.

git clone https://github.com/PapathanasiouIoannis/ML_Classification_NS.git
cd ML_Classification_NS/Bachelor_Thesis
python -m venv .venv
# Activate .venv using the command appropriate for your shell.
python -m pip install numpy pandas scipy sympy scikit-learn matplotlib seaborn joblib tqdm
python main.py

The first run can be computationally expensive: main.py uses all available CPU cores by default and requests 20,000 synthetic curves. Review TOTAL_CURVES, CURVES_PER_BATCH, and N_JOBS in main.py before running on a shared or resource-constrained system. Generated data are cached at Bachelor_Thesis/data/thesis_dataset.csv when the pipeline succeeds.

Relationship to later work

This historical repository should remain distinct from:

  • the post-thesis redevelopment, which contains the controlling methodological audit and more explicit provenance controls; and
  • the public interactive demonstrations, which are independent post-thesis extensions exposing retained artifacts under a deliberately narrow interpretation boundary.

Responsibility and supervision

Ioannis Papathanasiou performed the project work under the supervision of Charalampos Moustakidis and Theodoros Diakonidis at the Department of Physics, Aristotle University of Thessaloniki. Codex assisted with software development and documentation; scientific interpretation and responsibility remain with Ioannis Papathanasiou.

No software licence is currently included. Repository visibility alone should not be interpreted as permission to reuse the code.

About

Historical thesis-era code for a restricted synthetic hadronic-star versus CFL strange-star model comparison; not observational composition inference.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages