A structured collection of hands-on Machine Learning notebooks covering regression, classification, clustering, dimensionality reduction, and ensemble learning using Python and Scikit-learn.
This repository is focused on learning Machine Learning through practical implementation, visualization, experimentation, and model evaluation.
This repository contains topic-wise Jupyter Notebooks for learning and implementing Machine Learning algorithms from fundamentals to more advanced techniques.
Each topic is organized into its own folder and includes:
- 📓 Jupyter Notebook
- 📖 Topic-specific README
- 💻 Practical implementations
- 📊 Visualizations
- 📈 Model evaluation
- 🧠 Concept explanations
- 🔬 Hands-on experiments
The repository will be continuously updated as new Machine Learning concepts are learned and implemented.
| # | Algorithm / Topic | Files |
|---|---|---|
| 01 | Linear Regression | 1 |
| 02 | Multiple Linear Regression | 1 |
| 03 | Polynomial Regression | 1 |
| 04 | Logistic Regression | 2 |
| 05 | K-Nearest Neighbors | 1 |
| 06 | Decision Tree | 0 |
| 07 | Random Forest | 0 |
| 08 | Support Vector Machine | 0 |
| 09 | Naive Bayes | 0 |
| 10 | K-Means Clustering | 0 |
| 11 | Hierarchical Clustering | 0 |
| 12 | PCA | 0 |
| 13 | Gradient Boosting | 0 |
| 14 | XGBoost | 0 |
| 15 | DBSCAN | 1 |
| Total | 7 |
Progress is updated as new notebooks are added to the repository.
Machine-Learning-Notebooks/
│
├── 01_Linear_Regression/
│ ├── Linear_Regression.ipynb
│ └── README.md
│
├── 02_Multiple_Linear_Regression/
│ ├── Multiple_Linear_Regression.ipynb
│ └── README.md
│
├── 03_Polynomial_Regression/
│ ├── Polynomial_Regression.ipynb
│ └── README.md
│
├── 04_Logistic_Regression/
│ ├── Logistic_Regression.ipynb
│ └── README.md
│
├── 05_K_Nearest_Neighbors/
│ ├── KNN.ipynb
│ └── README.md
│
├── 06_Decision_Tree/
│ ├── Decision_Tree.ipynb
│ └── README.md
│
├── 07_Random_Forest/
│ ├── Random_Forest.ipynb
│ └── README.md
│
├── 08_Support_Vector_Machine/
│ ├── SVM.ipynb
│ └── README.md
│
├── 09_Naive_Bayes/
│ ├── Naive_Bayes.ipynb
│ └── README.md
│
├── 10_K_Means_Clustering/
│ ├── K_Means.ipynb
│ └── README.md
│
├── 11_Hierarchical_Clustering/
│ ├── Hierarchical_Clustering.ipynb
│ └── README.md
│
├── 12_PCA/
│ ├── PCA.ipynb
│ └── README.md
│
├── 13_Gradient_Boosting/
│ ├── Gradient_Boosting.ipynb
│ └── README.md
│
├── 14_XGBoost/
│ ├── XGBoost.ipynb
│ └── README.md
|
├── 15_DBSCAN/
│ ├── DBSCAN.ipynb
│ └── README.md
│
├── datasets/
│ └── README.md
│
├── README.md
├── requirements.txt
└── .gitignore
Regression algorithms focus on predicting continuous numerical values.
- Linear Regression
- Multiple Linear Regression
- Polynomial Regression
Classification algorithms predict discrete class labels.
- Logistic Regression
- K-Nearest Neighbors
- Decision Tree
- Random Forest
- Support Vector Machine
- Naive Bayes
Unsupervised learning techniques for discovering patterns and groups in data.
- K-Means Clustering
- Hierarchical Clustering
Techniques for reducing the number of features while retaining important information.
- Principal Component Analysis (PCA)
Methods that combine multiple models to improve predictive performance.
- Gradient Boosting
- XGBoost
- Random Forest
Each notebook follows a practical learning workflow:
Concept
↓
Theory
↓
Dataset
↓
Data Exploration
↓
Data Preprocessing
↓
Visualization
↓
Model Training
↓
Prediction
↓
Evaluation
↓
Experimentation
↓
Observations
The focus is on understanding both:
How the algorithm works
and
How to implement it in practice
| Category | Technologies |
|---|---|
| Programming Language | Python |
| Notebook Environment | Jupyter Notebook |
| Data Manipulation | NumPy, Pandas |
| Visualization | Matplotlib, Seaborn |
| Machine Learning | Scikit-learn |
| Gradient Boosting | XGBoost |
| Imbalanced Learning | Imbalanced-learn |
| Development Tools | Jupyter, VS Code |
| Version Control | Git, GitHub |
Additional libraries may be introduced as the repository expands.
git clone https://github.com/Aarush005coder/Machine-Learning-Notebooks.gitcd Machine-Learning-Notebookspython -m venv venvActivate it:
venv\Scripts\activatepython3 -m venv venvActivate it:
source venv/bin/activatepip install -r requirements.txtStart Jupyter Notebook:
jupyter notebookNavigate to the desired topic folder and open the corresponding .ipynb file.
For example:
04_Logistic_Regression/
└── Logistic_Regression.ipynb
Run the cells sequentially to reproduce the experiments, outputs, and visualizations.
Different algorithms require different evaluation metrics.
- MAE — Mean Absolute Error
- MSE — Mean Squared Error
- RMSE — Root Mean Squared Error
- R² Score — Coefficient of Determination
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- ROC Curve
- ROC-AUC
- Inertia
- Silhouette Score
The appropriate evaluation metric depends on the specific Machine Learning problem and dataset.
The notebooks use visualizations to make Machine Learning concepts easier to understand.
Examples include:
- Scatter plots
- Regression lines
- Distribution plots
- Correlation heatmaps
- Confusion matrices
- ROC curves
- Decision boundaries
- Clustering visualizations
- Dimensionality reduction plots
- Model performance visualizations
Notebook outputs and graphs are preserved in the .ipynb files whenever applicable, allowing results to be viewed directly on GitHub.
Each algorithm has its own dedicated folder:
Algorithm/
│
├── Algorithm.ipynb
└── README.md
If an algorithm requires multiple notebooks in the future, the naming convention will be:
Algorithm_01.ipynb
Algorithm_02.ipynb
Algorithm_03.ipynb
For example:
Logistic_Regression/
│
├── Logistic_Regression_01_Basics.ipynb
├── Logistic_Regression_02_Imbalanced_Data.ipynb
└── README.md
By working through this repository, you will build practical knowledge of:
- Machine Learning fundamentals
- Regression
- Classification
- Clustering
- Dimensionality reduction
- Ensemble learning
- Data preprocessing
- Exploratory Data Analysis
- Feature engineering
- Model training
- Model evaluation
- Hyperparameter tuning
- Data visualization
- Model interpretation
Machine Learning algorithms covered in this repository can be applied to problems such as:
- 🏠 House Price Prediction
- 💳 Credit Risk Analysis
- 🛡️ Fraud Detection
- 📧 Spam Detection
- 👥 Customer Segmentation
- 📉 Customer Churn Prediction
- 🏥 Medical Classification
- 📈 Sales Prediction
- 🛍️ Recommendation Systems
- 🔍 Anomaly Detection
The appropriate algorithm and evaluation strategy depend on the problem, dataset, and application requirements.
After creating or modifying a notebook:
git statusAdd the changes:
git add .Commit the changes:
git commit -m "Add new machine learning notebook"Push to GitHub:
git pushFor future updates, the same workflow can be used.
The long-term goal of this repository is to build a structured Machine Learning reference containing:
- 📖 Concepts
- 💻 Implementations
- 📊 Visualizations
- 🧪 Experiments
- 📈 Evaluation
- 🌍 Practical examples
The repository will continue to grow as new Machine Learning algorithms and concepts are explored.
This is primarily a personal Machine Learning learning repository.
Suggestions, corrections, and improvements are welcome.
If you find an issue:
- Open an issue.
- Describe the problem clearly.
- Suggest an improvement if possible.
Aarush Khandelwal
Artificial Intelligence & Data Science
- 🤖 Machine Learning
- 🧠 Artificial Intelligence
- 📊 Data Science
- 🔬 Deep Learning
- 🧩 Generative AI
- 💻 Data Structures & Algorithms
Learn the concept → Implement it → Visualize it → Evaluate it → Experiment with it.
Building Machine Learning knowledge one notebook at a time. 🚀