A local data engineering pipeline that transforms public GitHub events into analytics-ready Delta Lake datasets and an interactive portfolio dashboard.
Live dashboard · Architecture · Pipeline guide · Roadmap
Click the preview to explore the interactive dashboard.
This project demonstrates a production-style batch data pipeline without hiding the engineering logic inside notebooks. It emphasizes incremental ingestion, raw-data lineage, deterministic transformations, malformed-record quarantine, auditable quality checks, and bounded Spark-to-Pandas collections.
The implementation deliberately stays small enough to run locally while also supporting a thin Databricks Free Edition demonstration with the same modules.
flowchart LR
GH[GH Archive] --> DL[Hourly downloader]
DL --> RAW[Compressed JSON]
RAW --> B[(Bronze Delta)]
B --> S[Silver normalization]
S --> V[(Valid events)]
S --> R[(Rejected events)]
V --> Q[Quality checks]
R --> Q
V --> G[(Gold metrics)]
G --> D[Plotly dashboard]
D --> P[GitHub Pages]
- Ingestion: controlled GH Archive hourly downloads with streaming and gzip validation.
- Bronze: replayable raw events with source and ingestion metadata.
- Silver: explicit schemas, deterministic deduplication, specialized event tables, and rejected-record lineage.
- Quality: required-field, uniqueness, and cross-layer reconciliation checks persisted in Delta Lake.
- Gold: repository, daily activity, and exact period-summary metrics.
- Dashboard: a bounded, static Plotly artifact published with GitHub Pages.
See Architecture for the design decisions and table contracts.
- Python 3.11
- Apache Spark 3.5.9 and Delta Lake 3.3.x
- Docker Compose and JupyterLab
- Pandas and Plotly
- pytest, Ruff, and GitHub Actions
Docker with the Compose plugin is required.
make build
make up
make pipeline DATE=2025-01-02 START_HOUR=3 END_HOUR=8The pipeline runs each stage sequentially and stops on the first failure:
download → Bronze → Silver → quality → Gold → dashboard
Open docs/index.html to inspect the generated dashboard. Existing archives
and previously committed Bronze source files are skipped, so the same range can
be safely submitted again.
To remove all downloaded archives and generated Delta tables while preserving the tracked data directory placeholder, run:
make clean-dataFor individual commands, Jupyter access, and environment details, see Getting started. For layer behavior and outputs, see the Pipeline guide.
With the Spark service running:
make lint
make testThe test suite includes a synthetic end-to-end archive that crosses Bronze,
Silver, quarantine, Quality, Gold, and dashboard generation without downloading
external data. GitHub Actions executes the same checks for every pull request
and every push to master.
The local Docker environment remains the primary development profile. A thin Databricks notebook reuses the production modules and writes demo data to Unity Catalog Volumes without duplicating transformation logic.
See the Databricks guide for setup and execution. Evidence from a real Free Edition run includes the Git-backed notebook, Unity Catalog storage, quality results, Gold metrics, and a dashboard.
- Downloads cover one hour or a same-day UTC range; there is no scheduler or multi-day ingestion manifest.
- Silver and Gold use batch snapshot rebuilds instead of incremental merges.
- The static dashboard loads Plotly from a CDN.
Deferred work is tracked in the Roadmap. Airflow, Kafka, PostgreSQL, and MinIO are intentionally outside the current scope.
- Getting started: local setup and commands
- Pipeline guide: layer behavior and generated datasets
- Architecture: technical decisions and contracts
- Databricks guide: Free Edition execution profile
- Roadmap: completed work and future considerations
Available under the MIT License.