A structured collection of hands-on Machine Learning notebooks covering regression, classification, clustering, dimensionality reduction, and ensemble learning using Python and Scikit-learn.
This repository is focused on learning Machine Learning through practical implementation, visualization, experimentation, and model evaluation.
This repository contains topic-wise Jupyter Notebooks for learning and implementing Machine Learning algorithms from fundamentals to more advanced techniques.
Each topic is organized into its own folder and includes:
- π Jupyter Notebook
- π Topic-specific README
- π» Practical implementations
- π Visualizations
- π Model evaluation
- π§ Concept explanations
- π¬ Hands-on experiments
The repository will be continuously updated as new Machine Learning concepts are learned and implemented.
| # | Algorithm / Topic | Files |
|---|---|---|
| 01 | Linear Regression | 1 |
| 02 | Multiple Linear Regression | 1 |
| 03 | Polynomial Regression | 1 |
| 04 | Logistic Regression | 2 |
| 05 | K-Nearest Neighbors | 1 |
| 06 | Decision Tree | 0 |
| 07 | Random Forest | 0 |
| 08 | Support Vector Machine | 0 |
| 09 | Naive Bayes | 0 |
| 10 | K-Means Clustering | 0 |
| 11 | Hierarchical Clustering | 0 |
| 12 | PCA | 0 |
| 13 | Gradient Boosting | 0 |
| 14 | XGBoost | 0 |
| 15 | DBSCAN | 1 |
| Total | 7 |
Progress is updated as new notebooks are added to the repository.
Machine-Learning-Notebooks/
β
βββ 01_Linear_Regression/
β βββ Linear_Regression.ipynb
β βββ README.md
β
βββ 02_Multiple_Linear_Regression/
β βββ Multiple_Linear_Regression.ipynb
β βββ README.md
β
βββ 03_Polynomial_Regression/
β βββ Polynomial_Regression.ipynb
β βββ README.md
β
βββ 04_Logistic_Regression/
β βββ Logistic_Regression.ipynb
β βββ README.md
β
βββ 05_K_Nearest_Neighbors/
β βββ KNN.ipynb
β βββ README.md
β
βββ 06_Decision_Tree/
β βββ Decision_Tree.ipynb
β βββ README.md
β
βββ 07_Random_Forest/
β βββ Random_Forest.ipynb
β βββ README.md
β
βββ 08_Support_Vector_Machine/
β βββ SVM.ipynb
β βββ README.md
β
βββ 09_Naive_Bayes/
β βββ Naive_Bayes.ipynb
β βββ README.md
β
βββ 10_K_Means_Clustering/
β βββ K_Means.ipynb
β βββ README.md
β
βββ 11_Hierarchical_Clustering/
β βββ Hierarchical_Clustering.ipynb
β βββ README.md
β
βββ 12_PCA/
β βββ PCA.ipynb
β βββ README.md
β
βββ 13_Gradient_Boosting/
β βββ Gradient_Boosting.ipynb
β βββ README.md
β
βββ 14_XGBoost/
β βββ XGBoost.ipynb
β βββ README.md
|
βββ 15_DBSCAN/
β βββ DBSCAN.ipynb
β βββ README.md
β
βββ datasets/
β βββ README.md
β
βββ README.md
βββ requirements.txt
βββ .gitignore
Regression algorithms focus on predicting continuous numerical values.
- Linear Regression
- Multiple Linear Regression
- Polynomial Regression
Classification algorithms predict discrete class labels.
- Logistic Regression
- K-Nearest Neighbors
- Decision Tree
- Random Forest
- Support Vector Machine
- Naive Bayes
Unsupervised learning techniques for discovering patterns and groups in data.
- K-Means Clustering
- Hierarchical Clustering
Techniques for reducing the number of features while retaining important information.
- Principal Component Analysis (PCA)
Methods that combine multiple models to improve predictive performance.
- Gradient Boosting
- XGBoost
- Random Forest
Each notebook follows a practical learning workflow:
Concept
β
Theory
β
Dataset
β
Data Exploration
β
Data Preprocessing
β
Visualization
β
Model Training
β
Prediction
β
Evaluation
β
Experimentation
β
Observations
The focus is on understanding both:
How the algorithm works
and
How to implement it in practice
| Category | Technologies |
|---|---|
| Programming Language | Python |
| Notebook Environment | Jupyter Notebook |
| Data Manipulation | NumPy, Pandas |
| Visualization | Matplotlib, Seaborn |
| Machine Learning | Scikit-learn |
| Gradient Boosting | XGBoost |
| Imbalanced Learning | Imbalanced-learn |
| Development Tools | Jupyter, VS Code |
| Version Control | Git, GitHub |
Additional libraries may be introduced as the repository expands.
git clone https://github.com/Aarush005coder/Machine-Learning-Notebooks.gitcd Machine-Learning-Notebookspython -m venv venvActivate it:
venv\Scripts\activatepython3 -m venv venvActivate it:
source venv/bin/activatepip install -r requirements.txtStart Jupyter Notebook:
jupyter notebookNavigate to the desired topic folder and open the corresponding .ipynb file.
For example:
04_Logistic_Regression/
βββ Logistic_Regression.ipynb
Run the cells sequentially to reproduce the experiments, outputs, and visualizations.
Different algorithms require different evaluation metrics.
- MAE β Mean Absolute Error
- MSE β Mean Squared Error
- RMSE β Root Mean Squared Error
- RΒ² Score β Coefficient of Determination
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- ROC Curve
- ROC-AUC
- Inertia
- Silhouette Score
The appropriate evaluation metric depends on the specific Machine Learning problem and dataset.
The notebooks use visualizations to make Machine Learning concepts easier to understand.
Examples include:
- Scatter plots
- Regression lines
- Distribution plots
- Correlation heatmaps
- Confusion matrices
- ROC curves
- Decision boundaries
- Clustering visualizations
- Dimensionality reduction plots
- Model performance visualizations
Notebook outputs and graphs are preserved in the .ipynb files whenever applicable, allowing results to be viewed directly on GitHub.
Each algorithm has its own dedicated folder:
Algorithm/
β
βββ Algorithm.ipynb
βββ README.md
If an algorithm requires multiple notebooks in the future, the naming convention will be:
Algorithm_01.ipynb
Algorithm_02.ipynb
Algorithm_03.ipynb
For example:
Logistic_Regression/
β
βββ Logistic_Regression_01_Basics.ipynb
βββ Logistic_Regression_02_Imbalanced_Data.ipynb
βββ README.md
By working through this repository, you will build practical knowledge of:
- Machine Learning fundamentals
- Regression
- Classification
- Clustering
- Dimensionality reduction
- Ensemble learning
- Data preprocessing
- Exploratory Data Analysis
- Feature engineering
- Model training
- Model evaluation
- Hyperparameter tuning
- Data visualization
- Model interpretation
Machine Learning algorithms covered in this repository can be applied to problems such as:
- π House Price Prediction
- π³ Credit Risk Analysis
- π‘οΈ Fraud Detection
- π§ Spam Detection
- π₯ Customer Segmentation
- π Customer Churn Prediction
- π₯ Medical Classification
- π Sales Prediction
- ποΈ Recommendation Systems
- π Anomaly Detection
The appropriate algorithm and evaluation strategy depend on the problem, dataset, and application requirements.
After creating or modifying a notebook:
git statusAdd the changes:
git add .Commit the changes:
git commit -m "Add new machine learning notebook"Push to GitHub:
git pushFor future updates, the same workflow can be used.
The long-term goal of this repository is to build a structured Machine Learning reference containing:
- π Concepts
- π» Implementations
- π Visualizations
- π§ͺ Experiments
- π Evaluation
- π Practical examples
The repository will continue to grow as new Machine Learning algorithms and concepts are explored.
This is primarily a personal Machine Learning learning repository.
Suggestions, corrections, and improvements are welcome.
If you find an issue:
- Open an issue.
- Describe the problem clearly.
- Suggest an improvement if possible.
Aarush Khandelwal
Artificial Intelligence & Data Science
- π€ Machine Learning
- π§ Artificial Intelligence
- π Data Science
- π¬ Deep Learning
- π§© Generative AI
- π» Data Structures & Algorithms
Learn the concept β Implement it β Visualize it β Evaluate it β Experiment with it.
Building Machine Learning knowledge one notebook at a time. π