Skip to content

About

A collection of hands-on Machine Learning notebooks covering data preprocessing, EDA, visualization, supervised & unsupervised learning, model training, evaluation, and real-world datasets using Python and Scikit-learn.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

Β 

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ€– Machine Learning Notebooks

A structured collection of hands-on Machine Learning notebooks covering regression, classification, clustering, dimensionality reduction, and ensemble learning using Python and Scikit-learn.

This repository is focused on learning Machine Learning through practical implementation, visualization, experimentation, and model evaluation.


πŸ“Œ About

This repository contains topic-wise Jupyter Notebooks for learning and implementing Machine Learning algorithms from fundamentals to more advanced techniques.

Each topic is organized into its own folder and includes:

  • πŸ““ Jupyter Notebook
  • πŸ“– Topic-specific README
  • πŸ’» Practical implementations
  • πŸ“Š Visualizations
  • πŸ“ˆ Model evaluation
  • 🧠 Concept explanations
  • πŸ”¬ Hands-on experiments

The repository will be continuously updated as new Machine Learning concepts are learned and implemented.


πŸ“Š Machine Learning Progress

# Algorithm / Topic Files
01 Linear Regression 1
02 Multiple Linear Regression 1
03 Polynomial Regression 1
04 Logistic Regression 2
05 K-Nearest Neighbors 1
06 Decision Tree 0
07 Random Forest 0
08 Support Vector Machine 0
09 Naive Bayes 0
10 K-Means Clustering 0
11 Hierarchical Clustering 0
12 PCA 0
13 Gradient Boosting 0
14 XGBoost 0
15 DBSCAN 1
Total 7

Progress is updated as new notebooks are added to the repository.


πŸ—‚οΈ Repository Structure

Machine-Learning-Notebooks/
β”‚
β”œβ”€β”€ 01_Linear_Regression/
β”‚   β”œβ”€β”€ Linear_Regression.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 02_Multiple_Linear_Regression/
β”‚   β”œβ”€β”€ Multiple_Linear_Regression.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 03_Polynomial_Regression/
β”‚   β”œβ”€β”€ Polynomial_Regression.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 04_Logistic_Regression/
β”‚   β”œβ”€β”€ Logistic_Regression.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 05_K_Nearest_Neighbors/
β”‚   β”œβ”€β”€ KNN.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 06_Decision_Tree/
β”‚   β”œβ”€β”€ Decision_Tree.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 07_Random_Forest/
β”‚   β”œβ”€β”€ Random_Forest.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 08_Support_Vector_Machine/
β”‚   β”œβ”€β”€ SVM.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 09_Naive_Bayes/
β”‚   β”œβ”€β”€ Naive_Bayes.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 10_K_Means_Clustering/
β”‚   β”œβ”€β”€ K_Means.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 11_Hierarchical_Clustering/
β”‚   β”œβ”€β”€ Hierarchical_Clustering.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 12_PCA/
β”‚   β”œβ”€β”€ PCA.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 13_Gradient_Boosting/
β”‚   β”œβ”€β”€ Gradient_Boosting.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ 14_XGBoost/
β”‚   β”œβ”€β”€ XGBoost.ipynb
β”‚   └── README.md
|
β”œβ”€β”€ 15_DBSCAN/
β”‚   β”œβ”€β”€ DBSCAN.ipynb
β”‚   └── README.md
β”‚
β”œβ”€β”€ datasets/
β”‚   └── README.md
β”‚
β”œβ”€β”€ README.md
β”œβ”€β”€ requirements.txt
└── .gitignore

πŸ“š Topics Covered

πŸ”΅ Regression

Regression algorithms focus on predicting continuous numerical values.

  • Linear Regression
  • Multiple Linear Regression
  • Polynomial Regression

🟒 Classification

Classification algorithms predict discrete class labels.

  • Logistic Regression
  • K-Nearest Neighbors
  • Decision Tree
  • Random Forest
  • Support Vector Machine
  • Naive Bayes

🟠 Clustering

Unsupervised learning techniques for discovering patterns and groups in data.

  • K-Means Clustering
  • Hierarchical Clustering

🟣 Dimensionality Reduction

Techniques for reducing the number of features while retaining important information.

  • Principal Component Analysis (PCA)

πŸ”΄ Ensemble Learning

Methods that combine multiple models to improve predictive performance.

  • Gradient Boosting
  • XGBoost
  • Random Forest

🧠 Learning Approach

Each notebook follows a practical learning workflow:

Concept
   ↓
Theory
   ↓
Dataset
   ↓
Data Exploration
   ↓
Data Preprocessing
   ↓
Visualization
   ↓
Model Training
   ↓
Prediction
   ↓
Evaluation
   ↓
Experimentation
   ↓
Observations

The focus is on understanding both:

How the algorithm works

and

How to implement it in practice


πŸ› οΈ Technologies & Libraries

Category Technologies
Programming Language Python
Notebook Environment Jupyter Notebook
Data Manipulation NumPy, Pandas
Visualization Matplotlib, Seaborn
Machine Learning Scikit-learn
Gradient Boosting XGBoost
Imbalanced Learning Imbalanced-learn
Development Tools Jupyter, VS Code
Version Control Git, GitHub

Additional libraries may be introduced as the repository expands.


πŸ“¦ Installation

1. Clone the Repository

git clone https://github.com/Aarush005coder/Machine-Learning-Notebooks.git

2. Navigate to the Repository

cd Machine-Learning-Notebooks

3. Create a Virtual Environment

Windows

python -m venv venv

Activate it:

venv\Scripts\activate

macOS / Linux

python3 -m venv venv

Activate it:

source venv/bin/activate

4. Install Dependencies

pip install -r requirements.txt

πŸš€ Running the Notebooks

Start Jupyter Notebook:

jupyter notebook

Navigate to the desired topic folder and open the corresponding .ipynb file.

For example:

04_Logistic_Regression/
└── Logistic_Regression.ipynb

Run the cells sequentially to reproduce the experiments, outputs, and visualizations.


πŸ“Š Model Evaluation

Different algorithms require different evaluation metrics.

Regression

  • MAE β€” Mean Absolute Error
  • MSE β€” Mean Squared Error
  • RMSE β€” Root Mean Squared Error
  • RΒ² Score β€” Coefficient of Determination

Classification

  • Accuracy
  • Precision
  • Recall
  • F1-Score
  • Confusion Matrix
  • ROC Curve
  • ROC-AUC

Clustering

  • Inertia
  • Silhouette Score

The appropriate evaluation metric depends on the specific Machine Learning problem and dataset.


πŸ“ˆ Visualizations

The notebooks use visualizations to make Machine Learning concepts easier to understand.

Examples include:

  • Scatter plots
  • Regression lines
  • Distribution plots
  • Correlation heatmaps
  • Confusion matrices
  • ROC curves
  • Decision boundaries
  • Clustering visualizations
  • Dimensionality reduction plots
  • Model performance visualizations

Notebook outputs and graphs are preserved in the .ipynb files whenever applicable, allowing results to be viewed directly on GitHub.


πŸ“ Notebook Organization

Each algorithm has its own dedicated folder:

Algorithm/
β”‚
β”œβ”€β”€ Algorithm.ipynb
└── README.md

If an algorithm requires multiple notebooks in the future, the naming convention will be:

Algorithm_01.ipynb
Algorithm_02.ipynb
Algorithm_03.ipynb

For example:

Logistic_Regression/
β”‚
β”œβ”€β”€ Logistic_Regression_01_Basics.ipynb
β”œβ”€β”€ Logistic_Regression_02_Imbalanced_Data.ipynb
└── README.md

🎯 Learning Objectives

By working through this repository, you will build practical knowledge of:

  • Machine Learning fundamentals
  • Regression
  • Classification
  • Clustering
  • Dimensionality reduction
  • Ensemble learning
  • Data preprocessing
  • Exploratory Data Analysis
  • Feature engineering
  • Model training
  • Model evaluation
  • Hyperparameter tuning
  • Data visualization
  • Model interpretation

🌍 Real-World Applications

Machine Learning algorithms covered in this repository can be applied to problems such as:

  • 🏠 House Price Prediction
  • πŸ’³ Credit Risk Analysis
  • πŸ›‘οΈ Fraud Detection
  • πŸ“§ Spam Detection
  • πŸ‘₯ Customer Segmentation
  • πŸ“‰ Customer Churn Prediction
  • πŸ₯ Medical Classification
  • πŸ“ˆ Sales Prediction
  • πŸ›οΈ Recommendation Systems
  • πŸ” Anomaly Detection

The appropriate algorithm and evaluation strategy depend on the problem, dataset, and application requirements.


πŸ”„ Git Workflow

After creating or modifying a notebook:

git status

Add the changes:

git add .

Commit the changes:

git commit -m "Add new machine learning notebook"

Push to GitHub:

git push

For future updates, the same workflow can be used.


πŸ“Œ Repository Goals

The long-term goal of this repository is to build a structured Machine Learning reference containing:

  • πŸ“– Concepts
  • πŸ’» Implementations
  • πŸ“Š Visualizations
  • πŸ§ͺ Experiments
  • πŸ“ˆ Evaluation
  • 🌍 Practical examples

The repository will continue to grow as new Machine Learning algorithms and concepts are explored.


🀝 Contributions

This is primarily a personal Machine Learning learning repository.

Suggestions, corrections, and improvements are welcome.

If you find an issue:

  1. Open an issue.
  2. Describe the problem clearly.
  3. Suggest an improvement if possible.

πŸ‘¨β€πŸ’» Author

Aarush Khandelwal

Artificial Intelligence & Data Science

Areas of Interest

  • πŸ€– Machine Learning
  • 🧠 Artificial Intelligence
  • πŸ“Š Data Science
  • πŸ”¬ Deep Learning
  • 🧩 Generative AI
  • πŸ’» Data Structures & Algorithms

⭐ Learning by Building

Learn the concept β†’ Implement it β†’ Visualize it β†’ Evaluate it β†’ Experiment with it.

Building Machine Learning knowledge one notebook at a time. πŸš€

About

A collection of hands-on Machine Learning notebooks covering data preprocessing, EDA, visualization, supervised & unsupervised learning, model training, evaluation, and real-world datasets using Python and Scikit-learn.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages