Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

12 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

EEADataLakehouse

Two main areas of concern and class domains for the EEA data lakehouse:

  • eea_datalakehouse.data_preparation — data acquisition, vocabulary acquisition, data exploration, and transformation to parquet. Each stage is an abstract base class; concrete per-dataset pipelines subclass them since the actual logic varies by dataset and data flow.
  • eea_datalakehouse.dds_ingestion — notebook-side client for the Dremio Document Service (DDS) Ingest API (DI-8.4/8.5). FolderIngest transfers a folder of data files to S3 (via DDS-issued presigned URLs only — never an S3 SDK or S3 credentials) and registers it as a Dremio table, coordinated entirely through the DDS REST API.

Install

Install directly from GitHub, pinned to a released tag:

pip install "git+https://github.com/eeadata/[email protected]"

Usage

from pathlib import Path

from eea_datalakehouse.data_preparation import DataValidationError, ParquetTransformer


class MyDatasetTransformer(ParquetTransformer):
    def to_parquet(self, records, destination: Path) -> Path:
        ...  # dataset-specific parquet write


transformer = MyDatasetTransformer(required_fields=["id", "value"])

raw_source = [{"id": 1, "value": "a"}, {"id": 2, "value": "b"}]
records = transformer.prepare(raw_source)

try:
    transformer.validate(records)
except DataValidationError as exc:
    print(f"Invalid data: {exc}")

Dremio credentials are read from the injected kernel env vars _DREMIO_USER / _DREMIO_PWD, and the service URL from DDS_BASE_URL. None of these are ever logged or printed.

from eea_datalakehouse.dds_ingestion import FolderIngest

outcome = FolderIngest(
    folder="./my_data",
    target_catalog_path="biodiversity.uploads",
    data_format="parquet",      # one of parquet | csv | json
    intent="read_only",         # or "editable"
    conflict_mode="fail",
    parallelism=4,               # concurrent uploads (default 4)
).run()

print(outcome.commit.table_path, outcome.commit.record_count)

run() performs begin → upload(all files) → commit. Re-running the same session resumes by skipping files the server reports as already uploaded.

  • prepare(source) normalizes an iterable of raw records into a list of plain dicts.
  • validate(records) checks that every record contains the configured required_fields, raising DataValidationError on the first invalid record.
  • to_parquet(records, destination) is where a concrete subclass writes its own dataset's schema out to parquet.

Layout

Path Purpose
data_preparation/acquisition.py DataAcquirer — fetch a dataset's raw source data
data_preparation/vocabulary.py VocabularyLoader — load the controlled vocabulary to validate against
data_preparation/exploration.py Explorer — inspect acquired data before transformation
data_preparation/transformation.py ParquetTransformer — prepare/validate/write to parquet
dds_ingestion/credentials.py env-var creds + redacted DremioCreds
dds_ingestion/models.py typed request/response models for the DDS ingest contract
dds_ingestion/client.py thin, unit-testable HTTP client (IngestClient)
dds_ingestion/progress.py tqdm progress bar with graceful fallback
dds_ingestion/folder.py FolderIngest orchestration (scan/parallel/resume)

Development

pip install -e ".[dev]"
pytest

Releasing a new version

  1. Update version in pyproject.toml (e.g. 0.2.0).

  2. Commit and push to main.

  3. The Release GitHub Actions workflow runs on every push to main: it reads the version from pyproject.toml, builds the package, and publishes a v0.2.0 GitHub Release (creating the matching tag automatically), with notes auto-generated from merged PRs since the previous release (categorized via .github/release.yml).

    Pushing to main without bumping the version re-publishes the release for the current version with the latest build artifacts.

Install in JupyterLab

Run this in a notebook cell, pinned to the release tag you want:

%pip install "git+https://github.com/eeadata/[email protected]"

Use the %pip magic rather than !pip — it installs into the kernel the notebook is actually running on, instead of whatever pip happens to be first on PATH. After installing, restart the kernel (Kernel > Restart Kernel...) so the import below picks up the newly installed package:

from eea_datalakehouse.data_preparation import ParquetTransformer

Alternatively, download the wheel attached to the GitHub Release page for that tag and install the local file instead of pulling from git:

%pip install /path/to/EEADataLakehouse-0.1.0-py3-none-any.whl

About

Will hold all the python libraries for the EEA Data Lake House

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages