Skip to content

Latest commit

 

History

30 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Fast and Reliable Malware Detection Using Machine Learning

A Python-based offline malware and URL classification tool that analyzes Windows Portable Executable (PE) features and URL text patterns using machine-learning models.

This is an academic and research project. It is not a replacement for antivirus or endpoint-security software.

Objectives

  • Classify executable files as benign or malicious using static PE-header attributes.
  • Detect potentially malicious URLs using lexical text features.
  • Compare multiple machine-learning models.
  • Provide an offline terminal-based scanning workflow.
  • Evaluate model performance using reproducible validation techniques.

Detection Methods

PE Header Detection

The PE scanner extracts static attributes from Windows executable files using the pefile library. The extracted features are passed to a trained machine-learning classifier.

The scanner does not execute the submitted executable during analysis.

URL Detection

The URL scanner converts URL text into numerical features using TF-IDF vectorization and classifies URLs as benign or malicious.

A whitelist may be used to avoid flagging trusted domains, depending on the project configuration.

Dataset

PE Header Dataset

URL Dataset

The URL classifier uses a custom labeled dataset containing benign and malicious URLs.

Before training or evaluating the models, verify:

  • Label definitions.
  • Train/test split.
  • Duplicate records.
  • Class balance.
  • Whether URLs from the same domain appear in both training and test data.

Machine-Learning Models

Detection task Models evaluated Selected model Reported result
PE-header classification KNN, SVM, CNN, Random Forest Random Forest 99.37% accuracy
URL classification Logistic Regression Logistic Regression 98.46% accuracy

The reported results depend on the dataset, preprocessing, feature selection, train/test split, and evaluation procedure. They should not be interpreted as production detection rates.

Results

PE Header Classifier

  • Model: Random Forest.
  • Accuracy: 99.37%.

URL Classifier

  • Model: Logistic Regression.
  • Accuracy: 98.46%.
  • Precision: 99.18%.
  • Recall: 96.25%.

For complete evaluation, review the confusion matrix and ROC curve instead of relying only on accuracy.

Features

  • Terminal interface with ASCII-art branding using pyfiglet.
  • Static PE-header feature extraction.
  • URL classification using TF-IDF vectorization.
  • Optional whitelist filtering for trusted URLs.
  • Pre-trained model loading with joblib.
  • Offline scanning workflow.
  • Saved evaluation visualizations.
  • Optional Docker-based execution.

Project Structure

Update the file and directory names below if they differ from your repository:

Fast-Reliable-Malware-Detection-Using-KNN-Algorithm/
├── Classifier/
│   ├── pe_model.pkl
│   └── url_model.pkl
├── Dataset/
│   ├── pe_header_data.csv
│   └── url_data.csv
├── Extract/
│   └── feature_extraction.py
├── ML Model/
│   ├── pe_model_training.ipynb
│   └── url_model_training.ipynb
├── Screenshots/
├── main.py
├── requirements.txt
├── Dockerfile
├── README.md
├── LICENSE
└── .gitignore

Requirements

  • Python 3.9 or later.
  • scikit-learn.
  • pefile.
  • joblib.
  • pyfiglet.
  • pandas.
  • numpy.
  • matplotlib.

Example requirements.txt:

scikit-learn
pefile
joblib
pyfiglet
pandas
numpy
matplotlib

Use dependency versions compatible with your Python version. Very old versions may not install correctly on newer Python releases.

Installation

Clone the repository:

git clone [https://github.com/Adarshmishra87/Fast-Reliable-Malware-Detection-Using-KNN-Algorithm.git](https://github.com/Adarshmishra87/Fast-Reliable-Malware-Detection-Using-KNN-Algorithm.git)
cd Fast-Reliable-Malware-Detection-Using-KNN-Algorithm

Create a virtual environment:

python -m venv venv

Activate it on Windows:

venv\Scripts\activate

Activate it on Linux or macOS:

source venv/bin/activate

Install the dependencies:

pip install -r requirements.txt

Run the Application

Start the terminal interface:

python main.py

The application provides options similar to:

1: PE Scanner
2: URL Scanner

Follow the prompts shown by the application.

PE Scanner

Use a local PE file path when prompted:

C:\path\to\sample.exe

URL Scanner

Enter a URL when prompted:

[https://example.com](https://example.com)

Only scan files and URLs that you are authorized to analyze.

Docker Usage

Build the image:

docker build -t malware-detector .

Run the container:

docker run --rm -it malware-detector

To mount a controlled sample directory on Windows PowerShell:

docker run --rm -it `
  -v "${PWD}\samples:/samples" `
  malware-detector

The Docker workflow does not automatically make malware analysis safe. Do not execute suspicious files, and do not mount sensitive host directories.

Evaluation Visuals

View evaluation plots

URL Detector ROC Curve

URL Detector ROC Curve

URL Confusion Matrix

URL Confusion Matrix

PE Header Correlation Matrix

PE Header Correlation Matrix

PE Header Confusion Matrix

PE Header Confusion Matrix

Security and Safety Notes

  • Do not execute unknown or suspicious files on your normal operating system.
  • Use an isolated virtual machine or sandbox for malware-related experimentation.
  • Keep the virtual machine separate from personal accounts, credentials, and sensitive files.
  • Disable shared folders and clipboard integration when analyzing untrusted samples.
  • Do not upload private or confidential samples to public services.
  • Do not treat a benign classification as proof that a file is safe.
  • Do not treat a malicious classification as a confirmed threat without further analysis.
  • Machine-learning models can produce false positives and false negatives.
  • Keep dependencies updated.
  • Load joblib and pickle-based model files only from trusted sources.

Limitations

  • The project relies on dataset quality and feature availability.
  • PE-header analysis does not capture all runtime behavior.
  • URL classification may fail for obfuscated, shortened, newly registered, or previously unseen URLs.
  • Reported accuracy may not generalize to current real-world malware or URLs.
  • The system does not replace sandboxing, antivirus software, YARA rules, threat intelligence, or manual investigation.
  • The whitelist can cause trusted-looking URLs to bypass classification, so it must be maintained carefully.
  • Reported metrics may change if the dataset, preprocessing, or train/test split changes.

Future Improvements

  • Add reproducible train/test split configuration.
  • Add cross-validation and class-balance reporting.
  • Add precision, recall, F1-score, and ROC-AUC reporting for every model.
  • Add explainability using feature importance or SHAP.
  • Add configurable classification thresholds.
  • Add automated unit tests for feature extraction and prediction.
  • Add a web interface for controlled internal use.
  • Add model versioning and metadata.
  • Add YARA and hash-based checks as complementary analysis methods.
  • Add sandbox integration for authorized research environments.
  • Add CI checks for tests and dependency installation.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Author

Adarsh Mishra

Repository

Fast and Reliable Malware Detection Using Machine Learning

About

Fast and reliable malware detection system using the K-Nearest Neighbors (KNN) algorithm to classify executable files as benign or malicious based on extracted features such as byte sequences, metadata, and PE-file attributes.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages