Skip to content

Repository files navigation

GitHub Activity Lakehouse

A local data engineering pipeline that transforms public GitHub events into analytics-ready Delta Lake datasets and an interactive portfolio dashboard.

Python PySpark Delta Lake Docker CI License: MIT

Live dashboard · Architecture · Pipeline guide · Roadmap

GitHub Activity Lakehouse dashboard preview

Click the preview to explore the interactive dashboard.

Why this project

This project demonstrates a production-style batch data pipeline without hiding the engineering logic inside notebooks. It emphasizes incremental ingestion, raw-data lineage, deterministic transformations, malformed-record quarantine, auditable quality checks, and bounded Spark-to-Pandas collections.

The implementation deliberately stays small enough to run locally while also supporting a thin Databricks Free Edition demonstration with the same modules.

Architecture

flowchart LR
    GH[GH Archive] --> DL[Hourly downloader]
    DL --> RAW[Compressed JSON]
    RAW --> B[(Bronze Delta)]
    B --> S[Silver normalization]
    S --> V[(Valid events)]
    S --> R[(Rejected events)]
    V --> Q[Quality checks]
    R --> Q
    V --> G[(Gold metrics)]
    G --> D[Plotly dashboard]
    D --> P[GitHub Pages]
Loading
  • Ingestion: controlled GH Archive hourly downloads with streaming and gzip validation.
  • Bronze: replayable raw events with source and ingestion metadata.
  • Silver: explicit schemas, deterministic deduplication, specialized event tables, and rejected-record lineage.
  • Quality: required-field, uniqueness, and cross-layer reconciliation checks persisted in Delta Lake.
  • Gold: repository, daily activity, and exact period-summary metrics.
  • Dashboard: a bounded, static Plotly artifact published with GitHub Pages.

See Architecture for the design decisions and table contracts.

Technology

  • Python 3.11
  • Apache Spark 3.5.9 and Delta Lake 3.3.x
  • Docker Compose and JupyterLab
  • Pandas and Plotly
  • pytest, Ruff, and GitHub Actions

Quick start

Docker with the Compose plugin is required.

make build
make up
make pipeline DATE=2025-01-02 START_HOUR=3 END_HOUR=8

The pipeline runs each stage sequentially and stops on the first failure:

download → Bronze → Silver → quality → Gold → dashboard

Open docs/index.html to inspect the generated dashboard. Existing archives and previously committed Bronze source files are skipped, so the same range can be safely submitted again.

To remove all downloaded archives and generated Delta tables while preserving the tracked data directory placeholder, run:

make clean-data

For individual commands, Jupyter access, and environment details, see Getting started. For layer behavior and outputs, see the Pipeline guide.

Testing

With the Spark service running:

make lint
make test

The test suite includes a synthetic end-to-end archive that crosses Bronze, Silver, quarantine, Quality, Gold, and dashboard generation without downloading external data. GitHub Actions executes the same checks for every pull request and every push to master.

Databricks Free Edition

The local Docker environment remains the primary development profile. A thin Databricks notebook reuses the production modules and writes demo data to Unity Catalog Volumes without duplicating transformation logic.

See the Databricks guide for setup and execution. Evidence from a real Free Edition run includes the Git-backed notebook, Unity Catalog storage, quality results, Gold metrics, and a dashboard.

Known limitations

  • Downloads cover one hour or a same-day UTC range; there is no scheduler or multi-day ingestion manifest.
  • Silver and Gold use batch snapshot rebuilds instead of incremental merges.
  • The static dashboard loads Plotly from a CDN.

Deferred work is tracked in the Roadmap. Airflow, Kafka, PostgreSQL, and MinIO are intentionally outside the current scope.

Documentation

License

Available under the MIT License.

About

Local Data Engineering Lakehouse using Apache Spark, Delta Lake and GH Archive data.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages