A Python-based offline malware and URL classification tool that analyzes Windows Portable Executable (PE) features and URL text patterns using machine-learning models.
This is an academic and research project. It is not a replacement for antivirus or endpoint-security software.
- Classify executable files as benign or malicious using static PE-header attributes.
- Detect potentially malicious URLs using lexical text features.
- Compare multiple machine-learning models.
- Provide an offline terminal-based scanning workflow.
- Evaluate model performance using reproducible validation techniques.
The PE scanner extracts static attributes from Windows executable files using the pefile library. The extracted features are passed to a trained machine-learning classifier.
The scanner does not execute the submitted executable during analysis.
The URL scanner converts URL text into numerical features using TF-IDF vectorization and classifies URLs as benign or malicious.
A whitelist may be used to avoid flagging trusted domains, depending on the project configuration.
The URL classifier uses a custom labeled dataset containing benign and malicious URLs.
Before training or evaluating the models, verify:
- Label definitions.
- Train/test split.
- Duplicate records.
- Class balance.
- Whether URLs from the same domain appear in both training and test data.
| Detection task | Models evaluated | Selected model | Reported result |
|---|---|---|---|
| PE-header classification | KNN, SVM, CNN, Random Forest | Random Forest | 99.37% accuracy |
| URL classification | Logistic Regression | Logistic Regression | 98.46% accuracy |
The reported results depend on the dataset, preprocessing, feature selection, train/test split, and evaluation procedure. They should not be interpreted as production detection rates.
- Model: Random Forest.
- Accuracy: 99.37%.
- Model: Logistic Regression.
- Accuracy: 98.46%.
- Precision: 99.18%.
- Recall: 96.25%.
For complete evaluation, review the confusion matrix and ROC curve instead of relying only on accuracy.
- Terminal interface with ASCII-art branding using
pyfiglet. - Static PE-header feature extraction.
- URL classification using TF-IDF vectorization.
- Optional whitelist filtering for trusted URLs.
- Pre-trained model loading with
joblib. - Offline scanning workflow.
- Saved evaluation visualizations.
- Optional Docker-based execution.
Update the file and directory names below if they differ from your repository:
Fast-Reliable-Malware-Detection-Using-KNN-Algorithm/
├── Classifier/
│ ├── pe_model.pkl
│ └── url_model.pkl
├── Dataset/
│ ├── pe_header_data.csv
│ └── url_data.csv
├── Extract/
│ └── feature_extraction.py
├── ML Model/
│ ├── pe_model_training.ipynb
│ └── url_model_training.ipynb
├── Screenshots/
├── main.py
├── requirements.txt
├── Dockerfile
├── README.md
├── LICENSE
└── .gitignore
- Python 3.9 or later.
scikit-learn.pefile.joblib.pyfiglet.pandas.numpy.matplotlib.
Example requirements.txt:
scikit-learn
pefile
joblib
pyfiglet
pandas
numpy
matplotlib
Use dependency versions compatible with your Python version. Very old versions may not install correctly on newer Python releases.
Clone the repository:
git clone [https://github.com/Adarshmishra87/Fast-Reliable-Malware-Detection-Using-KNN-Algorithm.git](https://github.com/Adarshmishra87/Fast-Reliable-Malware-Detection-Using-KNN-Algorithm.git)
cd Fast-Reliable-Malware-Detection-Using-KNN-AlgorithmCreate a virtual environment:
python -m venv venvActivate it on Windows:
venv\Scripts\activateActivate it on Linux or macOS:
source venv/bin/activateInstall the dependencies:
pip install -r requirements.txtStart the terminal interface:
python main.pyThe application provides options similar to:
1: PE Scanner
2: URL Scanner
Follow the prompts shown by the application.
Use a local PE file path when prompted:
C:\path\to\sample.exe
Enter a URL when prompted:
[https://example.com](https://example.com)
Only scan files and URLs that you are authorized to analyze.
Build the image:
docker build -t malware-detector .Run the container:
docker run --rm -it malware-detectorTo mount a controlled sample directory on Windows PowerShell:
docker run --rm -it `
-v "${PWD}\samples:/samples" `
malware-detectorThe Docker workflow does not automatically make malware analysis safe. Do not execute suspicious files, and do not mount sensitive host directories.
View evaluation plots
- Do not execute unknown or suspicious files on your normal operating system.
- Use an isolated virtual machine or sandbox for malware-related experimentation.
- Keep the virtual machine separate from personal accounts, credentials, and sensitive files.
- Disable shared folders and clipboard integration when analyzing untrusted samples.
- Do not upload private or confidential samples to public services.
- Do not treat a benign classification as proof that a file is safe.
- Do not treat a malicious classification as a confirmed threat without further analysis.
- Machine-learning models can produce false positives and false negatives.
- Keep dependencies updated.
- Load
jobliband pickle-based model files only from trusted sources.
- The project relies on dataset quality and feature availability.
- PE-header analysis does not capture all runtime behavior.
- URL classification may fail for obfuscated, shortened, newly registered, or previously unseen URLs.
- Reported accuracy may not generalize to current real-world malware or URLs.
- The system does not replace sandboxing, antivirus software, YARA rules, threat intelligence, or manual investigation.
- The whitelist can cause trusted-looking URLs to bypass classification, so it must be maintained carefully.
- Reported metrics may change if the dataset, preprocessing, or train/test split changes.
- Add reproducible train/test split configuration.
- Add cross-validation and class-balance reporting.
- Add precision, recall, F1-score, and ROC-AUC reporting for every model.
- Add explainability using feature importance or SHAP.
- Add configurable classification thresholds.
- Add automated unit tests for feature extraction and prediction.
- Add a web interface for controlled internal use.
- Add model versioning and metadata.
- Add YARA and hash-based checks as complementary analysis methods.
- Add sandbox integration for authorized research environments.
- Add CI checks for tests and dependency installation.
This project is licensed under the MIT License. See the LICENSE file for details.
Adarsh Mishra
- GitHub: Adarshmishra87
- LinkedIn: adarsh-mishra-4b5792319



