This project was developed for the UIDAI Data Hackathon 2026. The goal is to solve the problem of "Digital Exclusion" by identifying geographic PIN codes where Aadhaar services (updates/biometrics) are lagging despite high enrolment numbers. We call these areas "Service Deserts".
- Automated Pipeline: Uses Python's glob library to dynamically merge 10+ large-scale CSV datasets.
- Inclusivity Index: A custom-engineered feature that calculates the ratio of Update Activity vs. Enrolment Saturation.
- Priority Mapping: Identifies the Top 10 Districts requiring immediate intervention via Mobile Aadhaar Vans.
- Data Cleaning: Robust handling of naming inconsistencies and "Noise" states (like State '0').
- Language: Python 3.10
- Libraries: Pandas, NumPy, Matplotlib, Seaborn, Glob
- Ingestion: Merging Biometric, Demographic, and Enrolment streams.
- Preprocessing: Normalizing district/state names and handling missing values.
- Analysis: Pincode-level spatial aggregation to avoid "District Averaging" bias.
- Insights: Identifying bottom-decile interaction zones.
- High concentration of Service Deserts in Kerala, Himachal Pradesh, and West Bengal.
- Critical MBU (Mandatory Biometric Update) Gap identified in the 5-17 age group.
- Pathanamthitta (Kerala) and Kangra (HP) ranked as highest priority for mobile unit deployment.
/api_data_.../: Source datasets (Categorized)master_pipeline.py: The core merging and cleaning scriptanalysis_viz.py: Script for generating insights and graphsFinal_Report.pdf: Detailed project documentation
- Alok Kumar - Data Scientist & Lead Developer

